Eric Schmidt: AI agents will outpace AlphaFold in science

Eric Schmidt, Google's CEO from 2001 to 2011 and now co-founder of the philanthropic venture Schmidt Sciences, and Suhas Mahesh, who leads AI for Science work at Schmidt Sciences' AI Center, argue in an MIT Technology Review op-ed that the AI model most people point to as proof science is about to be solved, AlphaFold, is the wrong template for what comes next. They open by noting that predictions of science's end recur every few decades: physicist Albert Michelson wrote in 1903 that the facts of physical science had all been discovered, and Stephen Hawking predicted in the 1980s that theoretical physics might be finished by the century's close. The current version of that feeling followed the 2024 Nobel Prize in chemistry awarded in part to Demis Hassabis and John Jumper of Google DeepMind for AlphaFold, the neural network that predicts three dimensional protein structures. Hassabis and his team called it the template for how AI can accelerate all of science to digital speed, and a wave of startups building foundation models for biology, chemistry and materials discovery raised billions of dollars on that promise.
The authors argue this template does not generalize. AlphaFold's success depended on the Protein Data Bank, a set of roughly 170,000 experimentally validated protein structures that took 53 years of international scientific cooperation and, by a recent estimate, about $21 billion of experimental work to assemble. Efforts of that scale are hard to fund and coordinate, so replicating the same condition in other fields will take decades rather than years. Even where money and cooperation exist, they add, most experimental science cannot produce comparably consistent data: cell lines drift, chemicals carry trace contaminants, lab humidity changes. Protein crystallography, the technique behind the Protein Data Bank, is unusually reliable, so much so that over 25 Nobel Prizes have relied on it; most fields have no equivalent. A handful of fields, including weather forecasting, much of genomics and narrow areas of chemistry, do meet the conditions and may see AlphaFold style breakthroughs soon.
For everything else, the authors point to AI agents, reasoning systems given access to tools, as the nearer term path. Because agents digitally model the iterative, judgment driven way scientists actually work rather than applying one trained model to one problem, they need far less specialized data. As an example, the authors describe Google's AI Co-Scientist, announced in May: given a one page brief and the goal of explaining how antibiotic resistance spreads between bacterial species, the system split the work across sub-agents, one drafting hypotheses from the literature, one critiquing them like a peer reviewer, one running tournaments to rank candidates and one refining the winner. It concluded that resistance genes hitch rides on bacterial viruses that ferry them into new hosts, a hypothesis the authors say was correct and that matched a still unpublished, still in peer review paper from Imperial College London researchers who had reached the same conclusion after a decade of wet lab work.
The authors credit agents with three compounding effects: a structural fix for science's reproducibility crisis, since agents automatically log every step and create an exact, replicable record instead of relying on researchers to voluntarily share raw data and code; an amplification of institutional memory, since a lab's history accumulates in a central, standardized repository rather than in messy notebooks handed down between generations of students; and, what they call the most important impact, speed. An agent able to read a thousand papers in an hour, design 500 molecules and learn from failed tests by morning lowers the cost of running an experiment, which the authors argue changes what researchers are willing to try. They place agentic AI alongside a short historical list of tools that reshaped every field at once, including calculus, statistical inference, spectroscopy and the computer.
The piece also names the limits still standing in the way: agents remain liable to hallucinate, their judgment is inconsistent, and memory and input constraints cap how long they can run autonomously. The authors expect these technical barriers to fall over time rather than resolving how or when. Additional research for the piece came from Maya Levin, an associate and sciences lead in Eric Schmidt's office.
Key facts
- AlphaFold was trained on the Protein Data Bank, roughly 170,000 experimentally validated protein structures assembled over 53 years at an estimated $21 billion in experimental work.
- Protein crystallography, the technique behind that dataset, is unusually reliable, underpinning over 25 Nobel Prizes; the authors say most experimental science lacks an equivalent, consistent measurement method.
- Google's AI Co-Scientist, announced in May, used sub-agents to hypothesize that antibiotic resistance genes spread via bacterial viruses, a conclusion the authors call correct and say matched a decade of unpublished wet-lab work by Imperial College London researchers.
- The op-ed is co-authored by Eric Schmidt, Google's CEO from 2001 to 2011 and co-founder of Schmidt Sciences, and Suhas Mahesh, who leads AI for Science work at Schmidt Sciences' AI Center, with additional research from Maya Levin.
- The authors list current agent limits: hallucination, inconsistent judgment, and memory and input constraints on how long they can run without supervision.
Why it matters
The debate the authors are entering is which AI approach actually scales across science. AlphaFold is the model everyone cites as proof AI can crack open a field, but the authors argue its recipe, a huge, unusually consistent, decades-in-the-making dataset, is close to unrepeatable elsewhere on any near-term timeline. That reframes where acceleration is likely to come from: not new single-purpose models trained the AlphaFold way, but general-purpose agents that reason the way working scientists do, combining imperfect evidence, judgment and revision rather than waiting for a perfect dataset to exist.
Who it affects
The argument targets researchers and labs in fields without an AlphaFold-grade dataset, which the authors say is most of biology and chemistry outside a few pockets like weather forecasting, genomics and narrow chemistry niches. It also speaks to the foundation-model startups in biology, chemistry and materials discovery that raised billions of dollars betting on the AlphaFold template, and to funders like Schmidt Sciences itself, which backs AI-for-science work through the same AI Center that Suhas Mahesh leads.
How to use it
This is an opinion piece, not a product launch, so there is nothing to buy or install. The practical takeaway the authors offer is a working example: Google's AI Co-Scientist, which researchers point at a one-page problem brief and let it split the work across hypothesis-drafting, critique, ranking and refinement sub-agents. The authors frame the near-term opportunity for labs as adopting that agent-driven workflow, and note a side benefit: agents log every action automatically, which the authors say gives labs a byproduct researchers have long resisted producing themselves, an exact, replicable record of method.
How solid is it
The piece is an op-ed by people with a direct stake in the argument: Eric Schmidt co-founded Schmidt Sciences, and Suhas Mahesh runs its AI-for-Science effort, so the piece is advocacy for a direction their own organization funds, not an independent study. The evidence for the central agent example is a single case, the Co-Scientist's antibiotic-resistance hypothesis matching prior wet-lab work at Imperial College London; no accuracy figures, error rates, or benchmark results are given beyond that one match, and the piece does not disclose what financial or institutional ties, if any, connect Schmidt Sciences to Google DeepMind or Google's AI Co-Scientist project beyond the authors' stated roles.
Risks and caveats
The authors themselves flag that agents are still liable to hallucinate, that their judgment is not consistent, and that memory and input constraints cap how long they can run on a problem without supervision, calling these real challenges rather than solved ones. The comparison case also rests on a single hypothesis match rather than a track record, and the underlying Imperial College London paper was still in peer review at the time of writing, with no year given for either that work or the Co-Scientist's antibiotic-resistance run.
“The most important impact of agents will be speed.”
— Eric Schmidt and Suhas Mahesh, MIT Technology Review