Study warns AI benchmark evidence doesn't always add up to real-world claims

A single AI benchmark score rarely stands alone as proof of anything useful. To reach a consequential claim, evaluators typically pass the result through several further steps: generalizing it to cases beyond the test set, interpreting it as evidence of a broader capability, extrapolating it to new tasks, transporting it to another system or deployment site, and combining it with assumptions about human review and downstream consequences. Existing validity-centred approaches to AI evaluation already ask for evidence supporting each of these individual claims.

This paper identifies a problem that sits one level up from that: warranted individual links do not automatically make a warranted chain. The target of one study may not match the starting point of the next; the system, population, outcome, or conditions being measured can quietly change at the interface between two steps; and evidence that looks independent can actually be dependent, for instance when two studies share the same underlying data or model lineage.

The paper builds its argument around the idea of projectibility: whether a bounded extension from observed cases to unobserved ones is actually warranted. It draws this framing from Goodman's classic problem of rival extensions, the observation that the same data can support multiple, incompatible extrapolations depending on which properties are treated as the ones worth projecting. Argument-based validity then supplies the machinery for testing which of those rival extensions can actually be defended.

The paper's central claim is what it calls a non-composition principle: support for two adjacent projections in an evaluation chain only combines into support for the full chain when the endpoints and assumptions of those projections line up, and when dependence and uncertainty are explicitly carried through the join rather than silently dropped. Two adjacent steps can each be individually well supported and still fail to add up to the larger claim being made.

The paper illustrates this with a legal-research case in which benchmark evidence and a separate deployment study are each sound on their own terms, yet remain parallel, meaning they do not actually connect into a single warranted claim about the deployed system. A further reanalysis and simulation show why this matters in practice: aggregate stability, a headline number that looks steady overall, can erase exactly the distinctions between subgroups or conditions that a later, more specific projection would need in order to be valid.

From this, the paper proposes a projectibility audit, a diagnostic procedure aimed at surfacing unsupported joins in the arguments that run from a benchmark result to a claim about real-world use, so that evaluators can spot where a chain of otherwise valid evidence has been stitched together incorrectly.

Key facts

  • Benchmark claims typically pass through several inferential steps before reaching a real-world conclusion: generalization to further cases, capability interpretation, extrapolation to new tasks, transport to another system or site, and combination with assumptions about human review and downstream consequences.
  • The paper's core finding is that warranted individual links do not automatically make a warranted chain: the target of one study may not match the source of the next, and system, population, outcome, or conditions can shift at the interface between two steps.
  • Evidence that looks independent can actually be dependent, for example when separate-seeming studies share the same underlying data or model lineage.
  • It frames the problem through Goodman's problem of rival extensions and pairs it with argument-based validity, then proposes a 'non-composition principle': adjacent projections combine only when their endpoints and assumptions align and uncertainty is carried through rather than dropped.
  • A legal-research case study and a separate reanalysis plus simulation show that benchmark evidence and deployment evidence can each be individually sound yet remain unconnected, and that aggregate stability can hide the distinctions a later projection needs, motivating a proposed 'projectibility audit' for finding these unsupported joins.

Why it matters

AI evaluation increasingly relies on stitching together several different kinds of evidence, a benchmark score here, a deployment study there, into one larger claim about whether a system is capable or safe enough for a given use. This paper's warning is that even when every individual step in that chain is well supported on its own, the combined chain can still be unwarranted, because nobody checked whether the steps actually connect. That is a more subtle failure mode than a single bad benchmark, and one that is easy to miss precisely because each piece looks solid in isolation.

Who it affects

The audience is AI evaluators, benchmark designers, and organizations that lean on benchmark results to justify deployment decisions, along with researchers who combine results across multiple studies or benchmarks. The paper's reanalysis point, that aggregate stability can erase the distinctions a later projection needs, is a direct warning to anyone summarizing benchmark results into a single headline figure.

How to use it

The paper offers a method rather than a product: a projectibility audit for checking whether a chain of benchmark-to-deployment reasoning actually holds together. Applying it means asking, at each join in the chain, whether the target of the earlier study matches the starting point of the next one, whether the system, population, outcome, or conditions changed at that interface, and whether apparently independent evidence is secretly dependent through shared data or model lineage. Only when those checks pass, and uncertainty is carried through rather than dropped, does the paper's non-composition principle allow treating the combined chain as warranted.

How solid is it

This is a methodological and conceptual paper rather than a large-scale empirical study, so its case for the non-composition principle rests on argument plus two illustrations: a legal-research case study showing benchmark and deployment evidence that stay parallel rather than combining, and a separate reanalysis and simulation showing how aggregate stability can mask distinctions relevant to later projections. There are no author names, institutional affiliation, or replication numbers in the available text, so those details cannot be verified from this source alone.

Risks and caveats

The paper's own point cuts both ways: a chain of benchmark results that looks stable and convincing in aggregate may still be hiding an unwarranted join, which means readers should not treat a steady headline number as proof that the underlying reasoning is sound. At the same time, the paper is not arguing that benchmarks are generally unreliable, only that composing several valid pieces of evidence into one larger claim requires its own separate justification. The projectibility audit is described as a diagnostic for finding these gaps, not a guarantee that any given chain of evidence is correct once applied.

“Warranted links don't automatically make a warranted chain.”

— the paper