Entity tracking emerges in language models at just 410 million parameters

Researchers evaluated entity tracking, the ability to follow where things are and how they change across a narrative even when this is not stated explicitly, in both language models and 48 human participants. The study used naturalistic narratives, ordinary story-like texts, at multiple levels of complexity, rather than the artificial tasks the authors say prior evaluations relied on, tasks they describe as far removed from natural language comprehension and never compared against human performance.

In the human group, entity tracking degraded specifically with narrative complexity, not with narrative length: longer stories were not harder to follow, but more complex ones were.

In language models, human-level entity tracking was already present in a model with just 410 million parameters, well below the multi-billion-parameter, code-specialized models that prior work had identified as capable of it. Performance kept improving with scale, and larger, contemporary models went on to exceed human performance.

The authors conclude that entity tracking, a core component of language understanding, emerges at model scales far smaller than previously thought.

Key facts

  • The study tested entity tracking, following where things are and how they change across a story, in 48 human participants and in language models using naturalistic narratives at several complexity levels.
  • In humans, entity-tracking accuracy fell as narrative complexity rose, but not with narrative length alone.
  • A language model with just 410 million parameters already matched human-level entity tracking, far below the multi-billion-parameter, code-specialized models that prior research had linked to the skill.
  • Entity tracking kept improving as model size grew, and larger, contemporary models exceeded human performance.
  • The authors conclude entity tracking, a core component of language understanding, emerges at far smaller model scales than previously thought.

Why it matters

Entity tracking, keeping straight who has what, where things are, and how that changes as a narrative unfolds, is a core component of language understanding: a system that loses track of it will misread a story even while producing fluent sentences. Prior work had placed this capability only in multi-billion-parameter, code-specialized models. This study finds it already at human level in a language model with just 410 million parameters, and finds that larger, contemporary models go on to exceed human performance outright. That reframes entity tracking as a capability that can appear well before frontier scale, not one that requires it.

Who it affects

The source does not name a company, institution, or specific model, so there is no vendor or product to point to directly. The likely audience is methodological: teams building or selecting small and mid-sized language models, researchers designing language-understanding benchmarks, and cognitive scientists comparing machine and human discourse comprehension. Anyone who assumed that tracking entities through a story requires a very large model now has a concrete data point against that assumption.

How to use it

There is no released tool or product here, only a research finding. The practical implication is for evaluation and model-selection choices: entity tracking does not appear to require reaching for the largest available model, and benchmarks meant to measure language understanding should include naturalistic narratives across a range of complexity levels, not just artificial tasks, and should compare results against a human baseline rather than scoring models in isolation.

How solid is it

The comparison draws on a human sample of 48 participants and naturalistic narratives at multiple complexity levels, rather than the artificial tasks the authors say earlier entity-tracking evaluations relied on. Several details that would help judge the result are not in the available text: how the narratives were sourced or written and how complexity levels were defined, why complexity rather than length degrades human performance, which specific models were tested, and a citation for the prior work that had placed entity tracking only in multi-billion-parameter, code-specialized models. The 410-million-parameter finding itself is stated as a direct result rather than an estimate, but it is presented here only at the level of detail an abstract provides.

Risks and caveats

The text available here is the paper's abstract, not the full methodology. The claim that contemporary models far exceed human performance is not accompanied by a figure for the size of that gap, so it should be read as directional rather than quantified. The abstract itself does not name authors or institutions, so the claims cannot be traced to a specific research group from this text alone, and the prior work said to have placed entity tracking only in multi-billion-parameter, code-specialized models is likewise not identified.