Jul 30, 2026
AI

Google DeepMind researcher says LLMs can't jump to new science

Tom Zahavy argues language models can derive and optimize, but lack the embodied leap needed to invent new scientific axioms.

Wei-Lin Zhao

By Wei-Lin Zhao · AI Correspondent

· 4 min read

Google DeepMind researcher says LLMs can't jump to new science
Photo: The Decoder

Google DeepMind researcher Tom Zahavy argues in a position paper titled “LLMs can't jump” that language models are unlikely to trigger scientific revolutions because they lack the mechanism for inventing genuinely new theoretical foundations. The claim matters for AI labs and enterprise buyers because it draws a boundary around current model capabilities: strong pattern matching and formal reasoning do not amount to the kind of creative break that produced general relativity.

Zahavy frames the argument around a model of discovery that Albert Einstein described in a letter to Maurice Solovine. In that account, scientific progress begins with sensory experience, moves through an intuitive leap toward axioms, and then uses deduction to generate testable results. Axioms are the starting assumptions of a theory, accepted before proof and used to build the rest of the system.

Why can't language models make scientific breakthroughs?

Zahavy uses philosopher Charles Sanders Peirce’s split between deduction, induction and abduction to locate the limitation. Deduction applies rules to reach conclusions. Induction generalizes from observed patterns. Abduction proposes an explanation for something surprising.

His distinction is between routine abduction and what he calls “manipulative abduction.” A model may select a plausible explanation from options that already exist, such as matching symptoms to a known disease. Zahavy argues that scientific invention often requires a harder move: creating an explanatory cause that has no existing linguistic template. That is the jump he says today’s language models cannot make.

The paper does not dismiss current systems’ technical progress. Zahavy says induction and deduction are areas where AI is already strong or advancing quickly. He points to AlphaProof, Gemini and GPT-5 reaching gold-level performance on International Mathematical Olympiad problems. He also concedes that a language model could likely derive general relativity if it were handed Einstein’s assumptions in advance. The unresolved step is producing those assumptions in the first place.

The paper’s example is Einstein’s break from Newtonian physics. Zahavy argues that an optimization-based AI would have lacked a reason to discard the existing theory because Newtonian mechanics was not facing a broad empirical collapse. The small irregularity in Mercury’s orbit had been explained at the time by positing an unseen planet, Vulcan. Measurements later used to support Einstein’s theory, including Eddington’s observation of light deflection, came after the theory had already been formulated.

What role does embodiment play in Zahavy's argument?

Zahavy points to Einstein’s “happiest thought,” the imagined experience of a person in free fall who no longer feels gravity. From that embodied mental simulation, Einstein developed the equivalence between gravity and acceleration, later illustrated through the thought experiment of an accelerating elevator in space.

The paper draws a similar line to the story of Archimedes arriving at the buoyancy principle through the physical experience of water rising in a bath. Zahavy’s point is that some scientific concepts begin with sensorimotor grounding before they can be formalized in language.

Language models, in this account, operate closer to John Searle’s Chinese Room thought experiment: they manipulate symbols without the physical experience that gives those symbols meaning. Zahavy applies the same skepticism to systems that automate research workflows. He says Sakana’s AI Scientist recombines existing concepts, while DeepMind’s AlphaEvolve is effective at optimization but still depends on a clear error signal to reduce.

Could world models change the limit?

Zahavy sees a possible path in physically consistent world models, especially systems that allow agents to act inside simulations. He distinguishes those from video generators such as Veo, which he describes as predicting likely next frames rather than understanding why events occur.

Action-controllable models such as Genie could let agents run counterfactual experiments inside synthetic environments, including interventions that resemble thought experiments. Zahavy’s argument is that such systems may provide the feedback loop needed for manipulative abduction, though the paper presents this as a path forward rather than evidence that current AI has crossed the line.

This story draws on original reporting from The Decoder.

More from AI

All AI →