Swapping the gold reference moves a six-language PII benchmark score by up to 7.55 F1 points
A benchmark score compares a system output against a reference. The authors of this paper argue that methodological attention falls almost entirely on the first term, the system output, and they set out to measure the second: the reference.
They work with a six-language benchmark for personally identifiable information (PII). Its retained annotation record holds three things: two independent annotator labellings, the aggregate that was shipped as the gold reference, and a reviewer gold produced by independent expert re-annotation of a sample. With that record, the authors hold the scored output fixed and exchange the reference.
The score moves. For one output it changes by 4.95 F1 points, and for the other by 2.00 (95% CIs [3.23, 6.27] and [0.26, 3.37]). In the worst language the shift is 7.55 F1 points.
The comparison between two outputs moves too. The paired interaction between reference and system is +2.95 points (CI [+1.97, +3.87]). The authors report that it survives correction for multiple testing and that it changes one language's margin outright.
The cause, according to the authors, is an undocumented aggregation default. It usually kept one annotator when adjudication did not fire, so the shipped reference ended up as a partial copy of an output being scored.
The paper reports the resulting reference-sensitivity band, shows how to compute one from any retained annotation record, and argues that this quantity belongs beside the score itself.
Key facts
- On a six-language PII benchmark, exchanging the reference while holding the scored output fixed moves F1 by 4.95 points for one output and 2.00 for the other (95% CIs [3.23, 6.27] and [0.26, 3.37]).
- In the worst language the shift reaches 7.55 F1 points.
- The paired interaction between reference and system is +2.95 points (CI [+1.97, +3.87]); it survives multiple-testing correction and changes one language's margin outright.
- The authors attribute the effect to an undocumented aggregation default that usually kept one annotator when adjudication did not fire, making the shipped reference a partial copy of an output being scored.
- They propose reporting a reference-sensitivity band beside the score and show how to compute it from any retained annotation record.
Why it matters
Benchmark scores are used to say one system beats another. This paper shows that, on one multilingual PII benchmark, the reference alone can move a score by several F1 points, and can change one language's margin outright. The authors say attention usually goes to the system output, not the reference, so this term is rarely examined. Here the reference turned out to be partly a copy of an output being scored, which is the kind of flaw a headline score cannot reveal.
Who it affects
Anyone who reads or builds on benchmark scores built from human annotation: people who compare systems on a leaderboard, and people who design benchmarks and ship a single aggregated gold set. The study concerns a PII benchmark covering six languages, so multilingual evaluation is the direct case.
How to use it
The authors say they show how to compute a reference-sensitivity band from any retained annotation record. The recipe in the paper uses what the benchmark already keeps: independent annotator labellings, the aggregate shipped as gold, and, where available, a reviewer gold from expert re-annotation of a sample. Hold the scored output fixed, swap the reference, and measure how far the score moves. They argue the resulting quantity belongs beside the score.
How solid is it
The claims are stated with uncertainty attached: 95% confidence intervals for both score shifts and for the paired interaction, and the authors report that the interaction survives correction for multiple testing. The cause is the authors' own diagnosis. The wording that one annotator was "usually" kept is the source's, and no exact proportion is given. The abstract does not name the benchmark, the six languages or the two scored outputs, and it does not say which language had the 7.55 shift.
Risks and caveats
The findings come from one six-language PII benchmark, and the abstract does not say whether the same aggregation default affects other benchmarks. It does not state the size of the expert re-annotation sample or the number of annotators, and it does not say whether the benchmark's maintainers responded or corrected the default. The method also needs a retained annotation record, so a benchmark that ships only the aggregated gold cannot be checked this way.
“The cause is an undocumented aggregation default that usually kept one annotator when adjudication did not fire, making the shipped reference a partial copy of an output being scored.”
— From the paper's abstract, arXiv 2610.03825