Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents
The paper "Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents" discusses the challenge of verifying the factuality of LLM agents that use multiple tools and sources through the Model Context Protocol (MCP). Existing systems like RAGAS faithfulness, MiniCheck, AlignScore, and SummaC check if claims are supported by pooled evidence but don't identify specific source support. The authors introduce ProvenanceGuard, which achieves a score of 0.802, outperforming other methods in source-aware verification.
This paper introduces ProvenanceGuard, which, unlike existing methods that pool evidence, specifically identifies which MCP tool output supports each claim, achieving a 0.802 score.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
PublishedOffset at this time: UTC+0Sep 29, 2026, 13:07 UTC
IngestedOffset at this time: UTC+0Sep 29, 2026, 14:00 UTC
- Published
- Sep 29, 2026, 13:07
- Ingested
- Sep 29, 2026, 14:00
- Source type
- Official
- Tier
- First-party
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
Tool-using LLM agents no longer read from a single retrieved passage. Through the Model Context Protocol (MCP), an agent can call a search tool, inspect a structured patient or account record, query a database, and pull metadata, then weave all of it into one answer. That m aakes the usual question of factuality more subtle than it looks. Most of the systems built to check LLM answers, from RAGAS faithfulness to fine-grained checkers like MiniCheck, AlignScore, and SummaC, ask whether a claim is supported by the available evidence once that evidence has been pooled together. In their usual form, they do not tell us which MCP tool output supports each claim, or whether that is the source the answer names.
Our latest paper, ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents (read it on Hugging Face, or on arXiv in the meantime), targets that gap. The failure mode we care about is one we call cross-source conflation: a claim that is true somewhere in the evidence, but attributed to the wrong source. A source-blind verifier may pass it, because the fact does exist in the pool. A source-aware verifier should not.
The problem: supported somewhere is not the same as supported by the right source
Consider a customer support agent that answers, "According to the account record, this plan includes a 30-day refund window." The refund window may be perfectly real, but stated in a policy document, not in the account record the answer points to. Pool the two together and the claim looks supported. Keep them separate and the attribution is wrong, and in a data-sensitive setting a wrong attribution can be as damaging as a wrong fact. The same pattern shows up in a clinical agent, where a patient-specific medication detail taken from a patient-history tool becomes misleading the moment the answer presents it as a finding from the medical literature.
A claim can be supported by one MCP source while the answer attributes it to another. Source-blind scoring sees support in the pooled evidence and passes it; ProvenanceGuard separately checks whether the supporting source matches the one the answer states or implies. Source: paper Figure 1.
This is why faithfulness scores, useful as they are, are not enough for MCP agents. An answer carries provenance, sometimes explicitly ("according to the account record") and sometimes implicitly. ProvenanceGuard keeps that connection between claim and source available for inspection.
What ProvenanceGuard does
ProvenanceGuard is a post-generation verification layer that sits on top of a black-box MCP agent. It runs after an agent produces an answer, and never collapses the evidence into one anonymous context. Instead it carries the source identity all the way through the pipeline. It reads the captured MCP trace, including the tool outputs and their source IDs, without retraining the agent. Then it does five things in sequence: it breaks the answer into specific claims, finds the source most relevant to each one, checks whether that source actually supports it, compares the source with the one the answer names or implies, and finally emits both a per-claim source verdict and a global, answer-level allow or block decision.
The verification flow. Source identity is preserved through decomposition, routing, support scoring, attribution checking, and repair, rather than being pooled. Blocked answers can go through RARR-style repair and be re-verified. Source: paper Figure 2.
A few of the design choices are worth calling out. For the experiments in our paper, we used local models so the captured traces could be processed in a controlled, offline setup: MiniLM helps find the relevant source, a DeBERTa NLI verifier model checks whether that source supports the claim, and a local language model helps break answers into claims. The verifier also checks literal values closely: a number, date, or identifier absent from the source cannot pass merely because the sentence sounds plausible. A calibrated decision step combines these signals. If an answer is blocked, a RARR -style repair step can try a source-grounded revision or a safe fallback, which the verifier then checks again.
Those named models are the setup we evaluated, not a requirement of ProvenanceGuard. The same claim, source, and decision steps can be adapted to hosted models where a team prefers cloud services; a new setup would need its own testing and calibration. Our reported results come from the local configuration. Its conservative decision policy suits data-sensitive review, where getting the source right matters more than producing the fastest possible answer.
Results
We tested ProvenanceGuard on answers from a medical agent that had used patient records, research articles, and other tools. This gave us 281 real traces to study. Medicine is a useful test because a fact from a patient's record and a fact from general research cannot be treated as the same source. The method can also be used in other fields when an agent keeps a record of its tool outputs and source IDs. For the main test, human experts checked 361 claims from 40 answers set aside from the data used to develop the system.
The most direct result is this: experts said 139 claims should not pass, and ProvenanceGuard caught 138 of them. It let one through. It also held 67 claims that the experts considered supported, sending them for review or repair. This reflects the cautious setting we tested: it favors a second look at some supported claims over letting unsupported ones through. For claims with an identifiable source, it also picked the right source about 86% of the time in this test.
We ran four other support checkers on the same claims. ProvenanceGuard scored highest on the paper's measure of how well a system catches claims that should be blocked while avoiding unnecessary blocks. The other checkers in this comparison did not tell us which tool output supported each claim. ProvenanceGuard records that connection, so a reviewer can see the source checked for each claim and the decision it produced.
Verifier Reject/block F1 Emits claim-to-source ID
ProvenanceGuard (ours) 0.802 Yes
MiniCheck 0.783 No
RAGAS Faithfulness 0.758 No
AlignScore 0.662 No
SummaC-ZS 0.436 No
Binary support metrics on the same held-out claim packet. ProvenanceGuard matches or beats the source-blind baselines on blocking while also producing per-claim source verdicts. Source: paper abstract and Table III.
Checking claims when sources look similar
In a separate, harder test with several similar sources, ProvenanceGuard scored 0.846 F1 for deciding which claims to block, but identified the exact source correctly in 50.3% of claims. Telling similar sources apart remains an important area for improvement.