Clinician preference is a poor proxy for LLM clinical safety, study finds
A new paper, "Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety," tests whether the standard way of ranking AI models, having people pick which of two answers they prefer, actually tells you whether those answers are medically safe. The authors, Fay Elhassan, David Sasu, Alexandra Kulinkina, Lars Henning Klein and Mary-Anne Hartley, drew on MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform that collects blinded pairwise preferences alongside separate multi-criterion rubric ratings. On the rubric, clinicians score each answer on a discrete scale from -2 to +2, where negative scores mean the content is clinically unsafe or misleading. The dataset covers 26,804 pairwise judgments on outputs from 13 different LLMs, contributed by more than 736 clinicians across 28 or more countries. The paper does not name the 13 models on its abstract page. The central finding: clinician preference is a poor proxy for safety-critical performance. Models that clinicians pick as the better answer in head-to-head comparisons can still rack up substantial rates of clinically meaningful failures, defined as a rubric score of -1 or below, on dimensions such as Harmlessness and Accuracy. Those failures are not spread evenly. They cluster unevenly across medical specialties, producing what the authors call domain-specific "no-go zones" that stay hidden in aggregate rankings or single-number leaderboards; the paper does not name which specialties. The authors also dig into why preference and safety diverge, examining prompt length, refusal and escalation behavior, and how much of a clinician's preference vote is driven by safety-critical features versus surface-level ones such as tone or formatting. They report that a substantial fraction of preference votes carry no positive safety signal at all, and that feature decomposition shows surface-level characteristics explain slightly more of the variation in preference than safety-critical rubric differences do; no specific percentages are given for either finding. To address the gap, the authors propose a clinically adjusted preference ranking that combines raw pairwise preference with the rubric-derived safety feedback, and report that it produces a more safety-aware ordering of the 13 models than the standard Bradley-Terry preference-strength score alone. The paper argues that evaluation practice should separate preference from safety, report safety-critical failure rates directly rather than folding them into a single score, and build clinically grounded adjustments into how LLMs are ranked for clinical decision-making use.
Key facts
- 26,804 pairwise judgments on outputs from 13 LLMs, contributed by more than 736 clinicians across 28+ countries, via the MOOVE platform.
- Clinicians also rate answers on a separate rubric scored -2 to +2, where a score of -1 or below counts as a clinically meaningful failure.
- Models that rank highly on pairwise preference can still show substantial rates of failure on Harmlessness and Accuracy.
- Those failures cluster unevenly by medical specialty, creating "no-go zones" invisible in aggregate leaderboard rankings.
- The authors propose a clinically adjusted ranking that blends preference with rubric-derived safety scores, yielding a more safety-aware ordering than raw Bradley-Terry preference strength.
Why it matters
AI model leaderboards, including clinical ones, typically rank systems by how often human judges prefer one answer over another. This paper is a direct empirical test of that method in a safety-critical domain and finds it wanting: a model can win preference comparisons while still producing unsafe or misleading medical content at a meaningful rate. That is a structural gap between what preference-based rankings measure and what they are often assumed to measure.
Who it affects
The finding bears on anyone who builds or relies on LLM evaluation leaderboards for clinical or health-adjacent use, including the developers of the 13 unnamed models tested, the clinicians and institutions using MOOVE-style evaluation, and, downstream, any organization choosing an LLM for clinical decision support based on preference rankings.
How to use it
The paper is a research contribution rather than a shipped tool: it proposes a clinically adjusted preference ranking method that combines pairwise preference scores with rubric-derived safety ratings, and reports that this combined ranking better reflects safety than raw preference strength (measured via the Bradley-Terry model) alone. It is available as a 27-page arXiv preprint with 10 figures, under a CC BY 4.0 license, for others to build on or apply.
How solid is it
The dataset is large by the standards of clinician-annotated LLM evaluation: 26,804 pairwise judgments from over 736 clinicians spanning 28-plus countries, collected on a dedicated platform (MOOVE) that gathers both blinded pairwise preferences and independent multi-criterion rubric ratings on the same outputs, which is what allows the authors to compare the two measures directly. The abstract page does not list author affiliations, and it gives no quantified figures for how large the "substantial fraction" of safety-blind votes is or exactly how much more variance surface-level features explain versus safety-critical ones, so those two claims are qualitative as stated.
Risks and caveats
The paper does not name the 13 LLMs evaluated, the specific medical specialties that form the "no-go zones," or a timeline for anyone adopting the proposed clinically adjusted ranking, so how directly these findings map onto any specific commercial model or specialty is not stated. The arXiv identifier suggests an August 2026 submission, but the page's own submission history lists 25 May 2026 for the v1 version; the source gives no explanation for the discrepancy.
“Models ranking highly under pairwise preference can still exhibit substantial rates of clinically meaningful failures on dimensions such as Harmlessness and Accuracy.”
— Elhassan, Sasu, Kulinkina, Klein and Hartley, "Preferred, Not Safer"