Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents

Published
Sep 29, 2026 — 13:07 UTC

Multiverse Computing has developed ProvenanceGuard, a source-aware factuality verification system designed for multi-contextual prompt (MCP) based large language model (LLM) agents. This system aims to address the issue of cross-source conflation, where claims may be true in some contexts but are incorrectly attributed to the wrong source. ProvenanceGuard enhances the transparency of source connections, allowing users to see the source of each claim made by the agent.

The system utilizes MiniLM to identify relevant sources and employs DeBERTa as a natural language inference (NLI) verifier to assess the support of these sources for the claims made. ProvenanceGuard demonstrated a high efficacy in its testing, successfully catching 138 out of 139 claims that should not pass verification. In total, the system evaluated 361 claims across 40 answers, with 67 claims deemed supported by expert review.

ProvenanceGuard achieved an F1 score of 0.802, indicating robust performance in source identification and verification. For comparison, other models were evaluated, with MiniCheck achieving an F1 score of 0.783, RAGAS at 0.758, AlignScore at 0.662, and SummaC-ZS at 0.436. Notably, the system also performed well in a similar sources test, achieving an F1 score of 0.846 and a correct source identification rate of 50.3%.

In terms of operational efficiency, ProvenanceGuard incurs a time overhead of approximately half a second per answer for verification processes. The system has also resolved 173 blocked answers, with fallback text instances used in 144 of these cases. This capability enhances the overall user experience by providing more reliable and contextually accurate responses from MCP-based LLM agents.

ProvenanceGuard was presented at the Agentic AI Summit at UC Berkeley in 2026, and the paper detailing this research is available on platforms such as Hugging Face and arXiv.

Summarised from Hugging Face Blog's coverage by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: Hugging Face Blog