AI bias audits detect bias but disagree on model rankings
Emerging AI regulation already mandates bias audits of high-risk systems, and the resulting audit scores are beginning to be used to rank models against each other. Both uses assume that different audit tools measure the same thing well enough to be compared. A new study tests that assumption directly: the authors ran ten extrinsic bias-audit instruments over a shared panel of ten frontier models, all queried through one pooled inference gateway, first for occupational gender bias, then for age and socioeconomic status.
Detection works. Eight of the ten tools found bias with confidence intervals clear of zero. Two widely cited direct-probe benchmarks, by contrast, came back saturated, because frontier models now answer those direct questions neutrally instead of revealing bias on them.
Ranking does not work. Agreement between the ten tools on how the ten models rank against each other is statistically indistinguishable from chance: Kendall's W is 0.07, with p equal to 0.83. To check whether the model panel was simply too similar in ability for any ranking to stabilize, the authors added a positive control of six deliberately weaker models. Within-tool reliability did recover once the panel spanned real capability gaps, meaning a single tool became more consistent as the range of model quality widened, but cross-tool ranking agreement still did not recover. The authors read that result as evidence that the ten instruments measure different underlying constructs, not the same construct with more or less noise.
The direction of the detected bias even splits by audit format. Forced-choice decision tools mostly over-corrected: toward women in general, and toward working-class candidates in 273 of 278 hiring decisions. Free-generation and default-coreference tools, in contrast, stayed stereotype-congruent, reproducing the conventional stereotype rather than correcting away from it. The same detect-but-disagree pattern replicates when the audits target socioeconomic status. For age, the tools appeared to agree on a ranking, but that apparent agreement dissolves once the paper's own rules for which tools to include are applied consistently.
The authors' practical message is that a single audit can detect bias and estimate its direction within its own operationalization, but no single audit supports ranking one model against another. All raw model responses, the code, and the analysis that recomputes every reported number from the source data are published on GitHub, so the results can be checked directly. The text does not name the paper's authors or institution, does not give a submission or publication date, and does not identify which ten specific audit instruments or which specific frontier models were tested, only their counts.
Key facts
- Ten extrinsic bias-audit instruments were run on a shared panel of ten frontier models through one pooled inference gateway, testing occupational gender bias, then age and socioeconomic status.
- Eight of the ten tools detected bias with confidence intervals clear of zero, while two widely cited direct-probe benchmarks came back saturated because frontier models now answer them neutrally.
- Cross-tool agreement on how the models rank against each other is statistically indistinguishable from chance: Kendall's W is 0.07, with p equal to 0.83.
- A positive control adding six deliberately weaker models showed within-tool reliability recovers once the panel spans real capability gaps, but cross-tool ranking still does not, which the authors read as the tools measuring different constructs.
- Bias direction splits by audit format: forced-choice tools mostly over-corrected, including toward working-class candidates in 273 of 278 hiring decisions, while free-generation and default-coreference tools stayed stereotype-congruent.
Why it matters
Emerging AI regulation already mandates bias audits of high-risk systems, and the resulting audit scores are starting to be used to rank models against each other, on the assumption that different tools are comparable. This study tests that assumption directly by running ten extrinsic bias-audit instruments over the same shared panel of ten frontier models, through one pooled inference gateway, across occupational gender bias, age and socioeconomic status. The result cuts against both regulatory uses at once: detection succeeds, since eight of the ten tools find bias with confidence intervals clear of zero, but ranking fails, since cross-tool agreement on which model is more or less biased is statistically indistinguishable from chance (Kendall's W is 0.07, p equal to 0.83). That gap matters for anyone treating a bias-audit score as a comparative metric: the authors' own practical message is that a single audit can detect bias and estimate its direction within its own operationalization, but no single audit supports ranking one model against another.
Who it affects
The finding is aimed at regulators and standards bodies writing bias-audit mandates for high-risk AI systems, since it questions whether different audit tools can be compared or combined into a ranking. It also affects any organization that procures or evaluates models partly on a published bias-audit score, and researchers who build, cite, or rely on the ten audit instruments tested here. The hiring-decision result, forced-choice tools over-correcting toward working-class candidates in 273 of 278 cases, points specifically at anyone using these audits to check or certify hiring-related AI systems. The text does not name the specific audit instruments, the ten frontier models, or the six added control models, so it identifies who is implicated only by count, not by name.
How to use it
There is no product, price, or access tier here: the output is a methodological finding plus a public reproduction package. All raw model responses, the code, and the analysis that recomputes every reported number from the source data are published in the paper's GitHub repository, so the results can be checked directly rather than taken on faith. The practical takeaway for anyone running or consuming a bias audit is to treat detection and ranking as two separate claims: a tool's finding that a model is biased, and in which direction, can be trusted within that tool's own setup, but its ranking of one model against another should not be read as a comparable, tool-independent measurement without corroboration from other instruments.
How solid is it
The design controls for an obvious confound by running all ten audit instruments on an identical panel of ten frontier models through a single pooled inference gateway, so differences between tools cannot be blamed on different prompting infrastructure. The key statistic, a cross-tool rank agreement of Kendall's W equal to 0.07 with p equal to 0.83, is about as close to pure chance as that measure gets. The authors also checked the most obvious alternative explanation, that the ten models were simply too similar in ability for any ranking to stabilize, by adding six deliberately weaker models as a positive control: within-tool reliability recovered once the panel spanned real capability gaps, but cross-tool ranking still did not, which argues against insufficient spread as the explanation. The same detect-but-disagree pattern held up when the authors reran the analysis for socioeconomic status, and an apparent cross-tool ranking agreement they found for age did not survive once the paper's own tool-inclusion rules were applied consistently, a self-check that cuts against their own result rather than for it. Against that strength, the text does not name the paper's authors, institution, or publication date, does not identify which specific instruments or models were used beyond their counts, and gives no mechanism for why forced-choice tools over-correct while free-generation and default-coreference tools stay stereotype-congruent, only that they do.
Risks and caveats
The central risk the paper documents is that a bias-audit score is already being used as if it were a portable ranking, and this result says that use is not statistically supported: cross-tool rank agreement across the ten instruments and ten models is indistinguishable from chance. A second risk is instrument saturation: two widely cited direct-probe benchmarks came back with no detectable signal because frontier models now answer them neutrally, even as eight of the ten instruments overall still detected bias with confidence intervals clear of zero, which leaves open the possibility that a model could read as bias-free on those two specific instruments while other instruments in the same study still flag it. A third is that bias direction itself is unstable across audit formats: forced-choice tools mostly over-corrected, including toward working-class candidates in 273 of 278 hiring decisions, while free-generation and default-coreference tools kept reproducing the conventional stereotype, so the same model can read as biased in opposite directions depending on which instrument is used. The source gives no percentage or rate for the 273-of-278 figure and no explanation for why audit format flips the bias direction, so both remain open questions the paper documents rather than resolves.
“a single audit can detect bias and estimate its direction within its own operationalization, but no single audit supports ranking one model against another.”
— the paper's abstract