AlphaGo veteran Thore Graepel argues LLMs don't reason

Thore Graepel, chair of machine learning at University College London and a former core member of the AlphaGo team at DeepMind, has written an opinion essay in MIT Technology Review with a blunt thesis: today's large language models do not reason, and the machinery that made AlphaGo's best-known move possible is still missing from them.
He starts in Seoul in March 2016, where he watched AlphaGo play move 37 in game two of its five-game match against Lee Sedol, one of the greatest professional Go players of all time. The stone went on the fifth line of the board and looked so absurd that some commentators thought it was a programming glitch. It wasn't. AlphaGo won the game and went on to win the match 4-1. Lee said afterwards: "I thought AlphaGo was based on probability calculation and that it was merely a machine. But when I saw this move, I changed my mind. Surely, AlphaGo is creative."
Graepel contrasts this with Deep Blue, which beat world chess champion Garry Kasparov in 1997 by looking six to eight moves ahead per player and evaluating 200 million positions per second, using rules hard-coded by humans. Go is far more complex: computing even a fraction of the possible outcomes would take a supercomputer billions of years. Many accounts portray move 37 as a flash of pure machine intuition. Graepel calls that a misunderstanding. AlphaGo has two systems. Its policy network, trained to guess what move a strong human would play, regarded move 37 as nothing special, a play with a roughly one in 10,000 chance of being made by an expert human. What made AlphaGo choose it was its search machinery, which looked beyond immediate plausibility and weighed the future consequences of proposed moves, explicitly building and searching a game tree with thousands of branches.
He maps this onto the System 1 / System 2 split popularized by Daniel Kahneman: System 1 is fast and gut-level, System 2 is slow, step-by-step and deliberative. AlphaGo's networks supplied the hunches and its search supplied the deliberation. Neither half works alone: intuition alone would never have chosen move 37, and brute-force search would have struggled to sift through all the possible moves.
A large language model, he says, picks the next token over and over, which amounts to System 1: fast, associative pattern completion. Chain of thought, where models generate intermediate steps before answering, was meant to fix this. The gains are real, above all in mathematics and coding, but Graepel says it does not add a genuinely separate reasoning mechanism. The intermediate reasoning is still produced by the same next-token prediction, iterated for longer before the model commits to an answer.
He lists three shortcomings that keep what chatbots do from qualifying as reasoning in a way a scientist might recognize. First, the models typically maintain no explicit, persistent, inspectable epistemic state: no open ledger of the hypotheses being considered, the confidence in various explanations, the evidence being weighed and the unresolved questions. Second, they lack a clean separation between what the system knows and how it manipulates that knowledge; both are interwoven in the network weights, with no independent, explicitly represented set of beliefs. Third, chains of thought look like deliberation, but research has demonstrated that the bots often concoct them after the fact, reaching an answer by one route and reporting another.
This matters, he argues, in high-stakes areas such as medicine, engineering and scientific research, where it matters not only what a system concludes but how. When a diagnosis or treatment goes wrong, people need to pinpoint whether the reasoning was at fault, the evidence was invalid or the assumptions were wrong.
He says this is why he recently left his position at Google DeepMind. He believes a fresh approach to machine reasoning is needed, one that draws on AlphaGo's architecture. AlphaGo keeps a record of what it knows about a position, the game tree, annotated with judgments from its neural networks, and updates it as reasoning progresses. A general reasoning system, he proposes, should keep an epistemic state of what it holds as settled, what it doubts, what it has ruled out and which questions stay open. Reasoning would then be a sequence of moves that change that state to reduce uncertainty: deducing consequences, breaking problems into parts, and deciding what question to ask, calculation to perform or experiment to run next.
He concedes that open-world reasoning is harder than Go or chess, since the state of affairs is only partially known, the set of actions is large and variable, and consequences are stochastic or unknown. But he says recent advances in LLMs and other neural models now make it possible to take this on: LLMs can suggest ways of tackling a problem given what is known and what resources are available, interact with tools via APIs or code, and help assess whether a claim is supported by evidence. To keep the system honest, an independent part must evaluate each move by how much it actually resolves uncertainty, updating beliefs only when the change is backed by evidence. Such a system could accumulate certified knowledge and improve its reasoning policy by learning from past reasoning, which he calls "the scientific method on steroids".
His conclusion: trustworthy machine intelligence will not come from making System 1 bigger. Scale sharpens intuition but does not make it more deliberative. He wants systems whose conclusions arise from an auditable sequence of evidence, inference and belief revision, in fields such as drug discovery, materials, climate and diagnosis, rather than from a convincing story told after the fact.
Key facts
- Thore Graepel, a former core member of the AlphaGo team at DeepMind, argues that AlphaGo's move 37 against Lee Sedol (game two, March 2016) came from its search, not pure intuition; its policy network gave the move roughly a one in 10,000 chance of being played by an expert human.
- He says an LLM picks the next token over and over, which amounts to System 1, and that chain of thought is still the same next-token prediction iterated for longer, not a separate reasoning mechanism.
- He names three shortcomings: no explicit, persistent, inspectable epistemic state; no clean separation between knowledge and how it is manipulated; and chains of thought that are often concocted after the fact.
- He proposes a system that keeps an epistemic state of what is settled, doubted, ruled out and open, with an independent part that updates beliefs only when backed by evidence, and says this is why he recently left Google DeepMind.
Why it matters
The essay offers a precise account of what is missing from current chatbots, from someone who helped build the system that beat a Go champion 4-1. Graepel says fluent output and even chain-of-thought gains in mathematics and coding do not amount to reasoning, because the same next-token prediction produces both the answer and the explanation. His claim is that trustworthy results and genuinely novel insights in science and medicine need a separate deliberative mechanism, as AlphaGo had, and that simply scaling up System 1 will not supply it.
Who it affects
Mainly people who rely on AI in high-stakes work. Graepel names medicine, engineering and scientific research, plus drug discovery, materials, climate and diagnosis, where he says it matters how a system reaches a conclusion and not only what it concludes. When a diagnosis or treatment goes wrong, he says, people need to tell whether the reasoning, the evidence or the assumptions failed. It also concerns researchers deciding whether to keep scaling LLMs or to pursue architectures that pair neural networks with explicit search and belief tracking.
How to use it
There is no product, model or tool to try. What the essay offers is a design idea. A reasoning system should hold an explicit epistemic state (settled, doubted, ruled out, open questions) and treat reasoning as moves that reduce uncertainty, such as deducing consequences, splitting problems into parts, or choosing the next question, calculation or experiment. LLMs would serve as components that propose approaches, call tools through APIs or code, and help judge whether a claim is supported by evidence, while an independent evaluator accepts belief updates only when evidence backs them. The essay does not present a working implementation, prototype or timeline for this architecture.
How solid is it
This is an opinion essay in MIT Technology Review, not a paper with experiments. The author's credentials are directly relevant: chair of machine learning at University College London and a core member of the AlphaGo team at DeepMind. The account of how AlphaGo's policy network and search divided the work is a first-hand description. The claims about LLMs are the author's argument. The essay does not name the research showing that chatbots concoct chains of thought after the fact and gives no figures for how often, and it gives no benchmark numbers on chain-of-thought gains beyond saying they are real, above all in mathematics and coding.
Risks and caveats
Graepel himself concedes that open-world reasoning is harder than Go or chess: the state of affairs is only partially known, the set of actions is large and variable, and consequences are stochastic or unknown. The proposed epistemic-state architecture is a plan, and the essay shows no working system. The author also has a stake in the argument, since he says the view is why he recently left Google DeepMind. The essay speaks of today's AI, large language models and chatbots in general and does not name specific models or companies as lacking reasoning, so how far each claim applies to any given product is left open.
“Scale sharpens intuition, but it does not make intuition more deliberative.”
— Thore Graepel, MIT Technology Review