SAGO paper measures LLM generalization by stability, not accuracy

The paper starts from a definition: generalization in large language models is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways.
The authors argue that existing work usually tests generalization through aggregate accuracy on a single prompt format, task, or set of variations. In their view this conflates robustness with overall benchmark performance. A single score can also be improved through narrow training or other means that obscure what generalization evaluation is supposed to show.
Their alternative looks at three things at once: individual examples, multiple input variants, and different aspects of model behavior. The focus is on variability rather than on reducing performance to one number.
From this view they introduce the Stability-Aware Generalization Objective (SAGO). It is a framework that measures how much model behavior changes for the same input under different variations and benchmarks. It captures variability across several dimensions, including generation consistency, internal activations, confidence, and response mirroring.
The headline finding is that many commonly used models exhibit statistically significant and consistent generalization instability. The authors summarize it in three parts: no model generalizes uniformly, the behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings. In other words, a model that leads on one dataset may not lead on another, and a model that is stable on one axis may be unstable on another.
Key facts
- SAGO, the Stability-Aware Generalization Objective, measures how much a model's behavior changes for the same input under different variations and benchmarks.
- It tracks several dimensions: generation consistency, internal activations, confidence, and response mirroring.
- The authors say aggregate accuracy on a single prompt format, task, or set of variations conflates robustness with overall benchmark performance.
- They report that many commonly used models show statistically significant and consistent generalization instability, and that no model generalizes uniformly.
- Behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings.
Why it matters
Most model comparisons rest on a benchmark score. The authors argue that such a score mixes up two different things: how good a model is on a benchmark and how robust it is when the same input is expressed differently. SAGO tries to pull them apart by looking at variability per example and per behavioral aspect instead of one aggregate number. The authors also say that a single score can be raised through narrow training, which hides the question of whether the model really generalizes. If their finding holds, a leaderboard position says little about how steady a model is.
Who it affects
The framing is aimed at anyone who evaluates or compares LLMs: researchers who build benchmarks and models, and teams who pick a model on the strength of published scores. The finding that cross-dataset variation can reverse model rankings bears directly on model selection, because a ranking from one dataset may not carry over to another.
How to use it
The source describes the framework and its findings but does not mention a release of code, data or a benchmark. What can be taken from it now is the evaluation idea: check the same example across several input variants, and look at more than one behavioral axis, namely generation consistency, internal activations, confidence and response mirroring, rather than relying on one accuracy figure.
How solid is it
The text is the paper's abstract as listed on Hugging Face Papers, so the claims are the authors' own and nothing here is independently checked. The abstract says the instability is statistically significant and consistent across many commonly used models. No numerical results, effect sizes, p-values or counts of models, datasets or prompt variants are given. The source does not say which models' rankings were reversed or on which datasets. No authors or institutions are named in the text.
Risks and caveats
The abstract says cross-dataset variation can reverse model rankings, not that it always does. It does not name the models that were tested or found unstable, so the finding cannot be tied to any specific product. The source also does not say whether SAGO yields a single score or a multi-axis profile beyond 'several dimensions'. Without the full paper, the size of the instability and the exact method remain unknown.
“no model generalizes uniformly, behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings”
— Abstract of the SAGO paper, Hugging Face Papers