DeepMind's Zahavy argues LLMs can't make the leap behind new science

DeepMind's Zahavy argues LLMs can't make the leap behind new science

A position paper from Google DeepMind researcher Tom Zahavy, titled 'LLMs can't jump,' argues that language models are structurally unable to spark a scientific revolution, no matter how capable they get at today's benchmarks. Zahavy builds his case on a framework Albert Einstein sketched in a letter to his friend Maurice Solovine: discovery is a cycle where sensory experience leads to an intuitive leap toward axioms, the unproven foundational assumptions of a theory, and from there logical deduction produces testable conclusions.

To locate the gap precisely, Zahavy uses philosopher Charles Sanders Peirce's three-way classification of reasoning. Deduction derives guaranteed conclusions from fixed rules, like a program producing a provably correct output. Induction spots patterns in data, for example generalizing from a thousand observed white swans that all swans are white. Abduction is the creative leap: inventing a cause to explain a surprising phenomenon. Zahavy concedes that language models already handle induction and deduction well. Systems such as AlphaProof, Gemini and GPT-5 now score at gold-medal level on International Mathematical Olympiad problems, and Zahavy even grants that a language model could probably derive general relativity if handed Einstein's assumptions as a starting point.

The bottleneck, he argues, sits inside abduction itself, which he splits into two levels. Ordinary abduction picks the most plausible explanation from a set of known candidates, the way a doctor matches symptoms to a disease, and language models can do this. 'Manipulative abduction' is harder: inventing a cause for which no linguistic template yet exists. That, Zahavy argues, is the actual bottleneck of scientific invention, and current machines cannot do it.

He illustrates the problem with Einstein's own case. AI systems typically learn by comparing predictions to reality and adjusting on the resulting error, but when Einstein was working there was no data crisis: Newtonian physics had been confirmed to extreme precision, and the only known anomaly, a small shift in Mercury's orbit, had already been explained away by positing a hypothetical hidden planet called 'Vulcan.' An optimization-driven system, Zahavy argues, would have had no error signal pushing it to overturn physics, and following that logic would have done what the era's astronomers did: invent an extra planet rather than rethink space and time. The observational evidence for Einstein's theory, such as Arthur Eddington's measurement of light deflection, did not arrive until years after the theory itself was formulated.

Zahavy traces Einstein's actual leap to what is known as his 'happiest thought': the freely falling observer who no longer feels gravity. That insight, Zahavy argues, came from embodied simulation, Einstein mentally running through a physical sensation rather than working through equations: he imagined a physicist inside an accelerating elevator in space and concluded that acceleration and gravity are indistinguishable from the inside. He draws a parallel to Archimedes, who by tradition discovered his buoyancy principle not through calculation but through the physical feeling of water rising as he stepped into a bathtub. In both cases a foundational principle emerged that had no name yet in the language of the time. Language models, Zahavy says, lack exactly this sensory grounding; he compares them to philosopher John Searle's 'Chinese Room' thought experiment, in which a person shuffles Chinese characters by rulebook without understanding any of them. Models manipulate the symbols of physics the same way, without the physical experience that gives those symbols meaning.

Zahavy also addresses two existing automated-science systems and finds both short of the leap. Sakana AI's 'AI Scientist' only recombines existing concepts, while DeepMind's own AlphaEvolve optimizes powerfully but needs a clear, shrinkable error signal, something Einstein never had. As a possible way past the gap, Zahavy points to action-controllable world models. He distinguishes these from video generators like Veo, which merely predict the statistically likeliest next frame, so an apple falls because falling is the most common continuation in training data rather than because the model understands gravity, something he still calls just pattern matching. Systems like Genie, by contrast, let an agent actively intervene in a simulation and run counterfactual experiments, such as mentally cutting an elevator's cable. A 'synthetic lab' built on that kind of model, Zahavy argues, could supply the feedback loop needed to invent genuinely new axioms where no linguistic template exists yet.

Key facts

  • Google DeepMind's Tom Zahavy, in a position paper titled 'LLMs can't jump,' argues language models lack the cognitive mechanism to spark a scientific revolution.
  • Using Charles Sanders Peirce's three-way split of reasoning, Zahavy says LLMs already handle deduction and induction well, citing gold-medal International Mathematical Olympiad scores from AlphaProof, Gemini and GPT-5, but he argues they cannot do 'manipulative abduction,' inventing a cause with no existing linguistic template.
  • He argues Einstein's leap to relativity had no error signal to optimize against, since Newtonian physics was confirmed to extreme precision at the time and the one anomaly, Mercury's orbit shift, had been explained by a hypothetical planet, 'Vulcan.'
  • Zahavy attributes Einstein's breakthrough to embodied simulation, the 'happiest thought' of a freely falling observer who no longer feels gravity, and draws a parallel to Archimedes' bathtub insight, arguing language models lack this sensory grounding, comparing them to Searle's 'Chinese Room.'
  • He finds Sakana AI's 'AI Scientist' and DeepMind's AlphaEvolve still short of the leap, and points to action-controllable world models like Genie, as opposed to next-frame predictors like Veo, as a possible path to genuine scientific abduction.

Why it matters

The paper is a direct challenge to the idea that scaling today's language models will eventually produce autonomous scientific discovery. Zahavy grants LLMs real strength at deduction and induction, gold-medal olympiad math among them, while arguing that the specific mechanism behind history's biggest theoretical leaps, manipulative abduction, is structurally out of reach for systems trained by minimizing prediction error against real data.

Who it affects

AI labs building toward autonomous or semi-autonomous scientific discovery, including DeepMind itself, whose own AlphaEvolve is named as falling short of the leap, and Sakana AI, whose 'AI Scientist' gets the same verdict. It also bears on anyone framing LLMs as future co-discoverers of new physics or similarly foundational science, since Zahavy's argument specifically targets that claim rather than narrower questions of coding or reasoning benchmarks.

How to use it

There is no product or release here; the paper functions as a framework for judging future systems. Its practical use is as a checklist: does a system merely predict the next likely outcome from training data, as Zahavy characterizes video generators like Veo, or can it actively intervene and run counterfactual experiments, as with action-controllable world models like Genie? Zahavy frames the latter as the more promising direction for building systems capable of manipulative abduction.

How solid is it

This is a position paper built on a philosophical framework, Peirce's categories of reasoning, and historical case studies, Einstein's relativity and Archimedes' buoyancy principle, rather than new experimental results or benchmark data showing LLMs failing at abduction directly. The strength of the argument rests on how well those historical analogies generalize, and on accepting Einstein's own retrospective account, via his letter to Solovine, of how his own thinking worked.

Risks and caveats

The argument leans on a small number of historical anecdotes, Einstein and Archimedes, to generalize about what any future scientific breakthrough requires, and both accounts come from tradition and retrospective self-report rather than direct observation of the reasoning process as it happened. The proposed fix, action-controllable world models such as Genie, is presented as a promising direction rather than a demonstrated solution: Zahavy argues it could supply the missing feedback loop, not that it already has.

“LLMs can't jump”

— title of Tom Zahavy's position paper