Stop calling AI's intermediate tokens 'reasoning traces,' a new paper argues
A position paper posted to arXiv argues that a common piece of AI vocabulary is doing real harm. Intermediate token generation, in which a language model produces a stretch of output before it reaches its final answer, has become a standard way to improve these models' performance on reasoning tasks. The paper's authors say the field's habit of labeling that intermediate output 'reasoning traces' or 'thinking traces' anthropomorphizes it: the term implies the tokens resemble the steps a human takes while working through a hard problem, and that reading them gives an end user an interpretable window into the model's actual thought process.
The authors argue this is not a harmless figure of speech. They say it is dangerous because it confuses what these models actually are and how to use them effectively, and that it has already led to what they call 'questionable research.' Their conclusion is a direct call to action: the community should stop anthropomorphizing intermediate tokens.
The abstract does not name or describe the specific 'questionable research' it references, and it gives no concrete examples of the anthropomorphizing language it is criticizing in papers, products or media coverage; it argues the general point rather than presenting new experiments or data. It also lists nine authors, led by Subbarao Kambhampati, who appears in the submission history as the arXiv submitter. The paper was first submitted on 14 April 2025 and has since been revised three times, most recently on 9 June 2026 (v4). On Hacker News, the submission drew 200 points and 122 comments.
Key facts
- The paper argues that calling a language model's intermediate output tokens 'reasoning traces' or 'thinking traces' anthropomorphizes them and is not a harmless metaphor.
- Intermediate token generation, producing output before the final answer, has become a standard technique for improving language models' performance on reasoning tasks.
- The authors say this anthropomorphized label confuses what these models are and how to use them effectively, and has already led to what they call 'questionable research.'
- The abstract lists nine authors, led by Subbarao Kambhampati, who appears in the submission history as the arXiv submitter.
- First submitted 14 April 2025 and revised through a fourth version on 9 June 2026, the paper drew 200 points and 122 comments on Hacker News.
Why it matters
Intermediate token generation, where a model produces a stretch of output before its final answer, has become a standard way to improve language models' performance on reasoning tasks. The paper's authors say the common label for that intermediate output, 'reasoning traces' or 'thinking traces,' anthropomorphizes it: it implies the tokens resemble the steps a human takes while solving a hard problem, and that reading them offers an end user an interpretable window into the model's actual thought process. They argue this framing is not a harmless figure of speech but a genuinely dangerous one, because it confuses what these models are and how to use them effectively.
Who it affects
The critique is aimed at anyone who builds, studies or relies on language models that produce this kind of output before an answer, including researchers who study these systems and end users who read a model's 'thinking' output as a faithful account of how it reached its answer. The submission drew engagement well beyond academia, gathering 200 points and 122 comments on Hacker News.
How to use it
The paper's practical ask is a change in vocabulary and mindset rather than a new tool: stop calling a model's intermediate output 'reasoning' or 'thinking' traces, and stop treating those tokens as though reading them gives a transparent view into the model's internal process. The authors present that mislabeling as a direct cause of confusion about how to use these models effectively, and as one of the sources behind what they call 'questionable research' built on that confusion.
How solid is it
This is a position paper, an argument rather than a new experiment or dataset: the abstract claims to present evidence for its central point but describes no additional empirical methodology of its own, and no publication venue or peer-review status is given. The version history shows the argument has been reworked repeatedly rather than settled quickly: first submitted 14 April 2025, and revised three times since, most recently on 9 June 2026 (v4).
Risks and caveats
The central claim that anthropomorphized language has caused 'questionable research' is something the abstract says it presents evidence for, rather than demonstrates with named specifics: the abstract does not name or describe the specific studies it has in mind, nor does it give concrete examples of the anthropomorphizing language it is criticizing in papers, products or media coverage. The argument rests on the authors' own interpretation of how the terminology is used and understood, not on cited evidence a reader can check against named cases.
“This anthropomorphization isn't a harmless metaphor, and instead is quite dangerous.”
— the paper's authors