Rationales barely raise QA accuracy but sway verifiers, DeepSeek tests show

Role-specialized question-answering pipelines increasingly pass a rationale from a reasoner to a verifier. The paper starts from the point that it is unclear what this message actually buys: better answers, stronger support assessment, or a new failure surface.

To find out, the authors introduce a message-intervention diagnostic. It fixes the evidence and the candidate answer and varies only the rationale passed across the reasoner-to-verifier boundary. The main experiment uses 400 examples drawn from MuSiQue, HotpotQA and 2WikiMultiHopQA, with DeepSeek as both generator and verifier.

The first result is that faithful rationales add almost no answer accuracy over giving no rationale at all. The second is that corrupted rationales strongly alter the verifier's support judgments. Under a blind verifier prompt, harmless paraphrases of the rationale shift support by only 0 to 2.5%, while corrupted rationales shift it by 10 to 22%. An explicit rationale-checking prompt amplifies the same pattern to 34 to 55%.

Final answers move less than support judgments, by 2 to 30%. Only 2.9 to 35.3% of corrupted support flips co-occur with an answer change, so the rationale can change what the verifier says about support without changing the answer.

Human audits show why this matters. Of the valid corruptions, 16 of 42 are corruption-overtrust cases. Blind humans reject or mark as unclear 9 of 10 audited corrupted rationales that the model accepts.

Cross-model and task-boundary checks show when the channel is active, amplified, inert, or folded into the task label. The paper concludes that rationale sharing should be evaluated as a verification-message mechanism, not merely as a route to higher answer accuracy.

Key facts

  • The diagnostic fixes the evidence and candidate answer and varies only the rationale passed from reasoner to verifier, tested on 400 examples from MuSiQue, HotpotQA and 2WikiMultiHopQA with DeepSeek as generator and verifier.
  • Faithful rationales add almost no answer accuracy over no rationale, while corrupted rationales strongly alter support judgments.
  • Under a blind verifier prompt, harmless paraphrases shift support by 0 to 2.5% and corrupted rationales by 10 to 22%; an explicit rationale-checking prompt raises the corrupted figure to 34 to 55%.
  • Final answers move 2 to 30%, and only 2.9 to 35.3% of corrupted support flips co-occur with answer changes.
  • In human audits, 16 of 42 valid corruptions are corruption-overtrust cases, and blind humans reject or mark unclear 9 of 10 audited corrupted rationales the model accepts.

Why it matters

Passing rationales between a reasoner and a verifier is becoming common in role-specialized QA pipelines, yet the paper says it is unclear what the message buys. Its answer reframes the question: on these tests the rationale does little for accuracy but a lot for the verifier's judgment of support. The authors argue that rationale sharing should be evaluated as a verification-message mechanism, not merely as a route to higher answer accuracy.

Who it affects

Anyone building or evaluating multi-agent or role-split QA systems in which one model explains its reasoning and another checks it. It also concerns evaluators who judge rationales mainly by whether final-answer accuracy improves, since the paper finds that measure misses most of the effect.

How to use it

The method itself is the practical takeaway. Hold the evidence and candidate answer constant, then vary only the rationale handed to the verifier: no rationale, a faithful one, a harmless paraphrase, a corrupted one. Compare support judgments as well as final answers. The paper also reports that the verifier prompt matters: a blind prompt and an explicit rationale-checking prompt gave different sizes of effect.

How solid is it

The main experiment is modest in scope: 400 examples across three multi-hop QA datasets, with DeepSeek as generator and verifier. The claims are the authors' own results, and a human audit supports the overtrust finding, though the audit of corrupted rationales is small (9 of 10 refers to 10 audited cases). The specific DeepSeek model or version is not stated, and it is not stated whether the percentage ranges are percentage points or relative changes, nor what the baselines are. Cross-model results are described only qualitatively.

Risks and caveats

The main risk the paper points to is corruption-overtrust: a verifier that accepts a flawed rationale that blind humans reject or find unclear. Because only 2.9 to 35.3% of corrupted support flips co-occur with answer changes, accuracy checks alone can hide this. The source does not say how corruptions were generated, and the ranges are not broken down by dataset. The cross-model and task-boundary checks suggest the effect is not uniform: the channel can be active, amplified, inert, or folded into the task label.

“Rationale sharing should be evaluated as a verification-message mechanism, not merely as a route to higher answer accuracy.”

— From the paper's abstract