Multiverse Computing's ProvenanceGuard checks the source behind MCP agent claims

Multiverse Computing has published a blog post on its paper ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents. The paper targets a failure mode the team calls cross-source conflation: a claim that is true somewhere in the evidence, but attributed to the wrong source. A source-blind verifier may pass it, because the fact does exist in the pool of evidence. A source-aware verifier should not.
The post gives two examples. A customer support agent says, "According to the account record, this plan includes a 30-day refund window." The refund window may be real, but stated in a policy document rather than the account record the answer points to. Pool the two and the claim looks supported; keep them apart and the attribution is wrong. In a clinical agent, a patient-specific medication detail taken from a patient-history tool becomes misleading once the answer presents it as a finding from the medical literature. The authors argue that faithfulness scores alone are not enough for MCP agents, since an answer carries provenance, sometimes explicitly and sometimes implicitly.
ProvenanceGuard is a post-generation verification layer that sits on top of a black-box MCP agent. It reads the captured MCP trace, including tool outputs and their source IDs, without retraining the agent, and never collapses the evidence into one anonymous context. It works in five steps: it breaks the answer into specific claims, finds the source most relevant to each, checks whether that source supports the claim, compares that source with the one the answer names or implies, and emits both a per-claim source verdict and a global, answer-level allow or block decision. The verifier checks literal values closely, so a number, date or identifier absent from the source cannot pass just because the sentence sounds plausible. A calibrated decision step combines the signals. If an answer is blocked, a RARR-style repair step can attempt a source-grounded revision or a safe fallback, and the verifier then checks it again.
In the paper's experiments the team used local models so the traces could be processed offline in a controlled setup: MiniLM to find the relevant source, a DeBERTa NLI verifier to check support, and a local language model to split answers into claims. The authors say these are the evaluated setup, not a requirement; the steps could be adapted to hosted models, but a new setup would need its own testing and calibration. The reported results come from the local configuration.
The main test used a medical agent that had drawn on patient records, research articles and other tools, giving 281 real traces. Human experts checked 361 claims from 40 answers set aside from the data used to develop the system. Experts said 139 claims should not pass, and ProvenanceGuard caught 138 of them, letting one through. It also held 67 claims the experts considered supported, sending them for review or repair; the authors describe this as the cautious setting, favouring a second look at some supported claims over letting unsupported ones through. For claims with an identifiable source, it picked the right source about 86% of the time. The team ran four other support checkers on the same claims. ProvenanceGuard scored highest on the paper's measure of catching claims that should be blocked while avoiding unnecessary blocks, and the other checkers did not say which tool output supported each claim.
In a separate, harder test with several similar sources, ProvenanceGuard scored 0.846 F1 for deciding which claims to block, but identified the exact source correctly in 50.3% of claims. The authors call telling similar sources apart an important area for improvement. In a controlled test of wrong attribution, they changed the named source in 50 cases while leaving the supporting evidence intact, and ProvenanceGuard caught all 50 swaps.
On repair, the full-trace run with the RARR-style loop resolved all 173 blocked answers, though 144 of them ended in fallback text rather than a substantive rewrite; the authors frame this as the system avoiding an unverifiable answer rather than manufacturing one. On reconstructed multi-source test traces, a fresh repair run resolved all 59 initially blocked answers with only two terminal fallbacks. As an offline gate the overhead is roughly half a second per answer on the reported local configuration, with the NLI and routing calls themselves in the tens of milliseconds.
The post says the approach has already been adapted in NVIDIA NVFlow, which merged an optional grounding-verification stage for its finance agent. That stage checks completed answers against the SEC excerpts the agent retrieved and saves separate decisions without changing the original rollout or training data. It uses ProvenanceGuard's source-aware verification approach; the repair loop belongs to the broader research system. ProvenanceGuard was also presented as a poster at the Agentic AI Summit 2026 at UC Berkeley.
Key facts
- ProvenanceGuard is a post-generation layer over a black-box MCP agent that reads the captured trace and source IDs, without retraining the agent, and flags claims attributed to the wrong source ("cross-source conflation").
- On a medical-agent test of 361 expert-checked claims from 40 held-out answers, it caught 138 of the 139 claims experts said should not pass, and held 67 claims the experts considered supported.
- In a harder test with several similar sources it scored 0.846 F1 for blocking, but identified the exact source in only 50.3% of claims; it caught all 50 deliberate attribution swaps in a controlled test.
- With a RARR-style repair loop, the full-trace run resolved all 173 blocked answers, but 144 of those ended in fallback text; overhead is roughly half a second per answer on the reported local setup.
- NVIDIA NVFlow merged an optional grounding-verification stage for its finance agent that uses the source-aware approach.
Why it matters
As agents move from single-passage RAG to multi-tool MCP setups, the authors argue that which source a fact came from stops being a footnote and becomes part of what factuality means. A checker that pools all evidence can pass a true statement credited to the wrong source, and in data-sensitive settings such as clinical or customer-account work a wrong attribution can be as damaging as a wrong fact. ProvenanceGuard keeps the claim-to-source link visible, so a reviewer can see which source was checked for each claim and what decision followed.
Who it affects
Teams building or auditing MCP-based agents that combine several tools with different kinds of records, such as patient history versus medical literature, or account records versus policy documents. It is aimed at data-sensitive review, where the local, offline configuration keeps sensitive traces in a controlled environment. The NVFlow finance-agent stage is an example of another group adapting the approach.
How to use it
The layer runs after the agent produces an answer and needs the captured MCP trace with tool outputs and source IDs; no retraining of the agent is required. The evaluated setup used MiniLM for source routing, a DeBERTa NLI verifier for support, and a local language model for claim decomposition. The authors say the steps can be adapted to hosted models, but a new setup would need its own testing and calibration. They also say the method can be used in other fields when an agent keeps a record of its tool outputs and source IDs. Blocked answers can go through a RARR-style repair and be re-verified.
How solid is it
The figures come from the authors' own blog post about their paper, and the main test is human-expert-checked on 361 claims from 40 held-out answers. Results are from a medical agent only; no test on other domains is reported by the authors, and the NVFlow adaptation has no reported metrics. No numeric scores are given for the four other checkers in the main comparison, and they are not named, so the claim that ProvenanceGuard scored highest rests on the paper's own measure. The main-test F1 of ProvenanceGuard is not stated in the post; only the 138 of 139 caught and 67 held are given.
Risks and caveats
The policy is deliberately conservative: 67 claims that experts considered supported were held for review or repair. Source identification is much weaker when sources look alike: 50.3% exact-source accuracy in the harder test, against about 86% in the main test. In the full-trace repair run, 144 of 173 resolved answers ended in fallback text rather than a substantive rewrite. The reported results come from the local configuration, so a hosted-model setup would need its own calibration. The text does not say the tool is released as open source or available as a product.
“a claim that is true somewhere in the evidence, but attributed to the wrong source”
— Multiverse Computing, describing cross-source conflation