ClaimReceipt Verifies Evidence for Agent Evaluations

Peiying Zhu, Sidi Chang· September 3, 2026 View original

Key takeaways

  • Reliable AI agent evaluation requires verifiable evidence sufficiency and experiment coverage.
  • ClaimReceipt is a new specification and verifier for binding evidence to experiment manifests.
  • It provides clear audit verdicts (PASS, INVALID, INCONCLUSIVE) for agent claims.
  • ClaimReceipt adds minimal overhead while significantly enhancing evaluation rigor and transparency.

Who benefits

AI/ML EngineeringComplianceAuditingResearch & DevelopmentCybersecurity

Summary

ClaimReceipt is a new specification and verifier designed to ensure evidence sufficiency and coverage in AI agent evaluations. It binds typed transaction evidence to a signed experiment manifest, providing reliable audit verdicts and addressing the limitations of generic logs for verifying agent claims.

Evaluating AI agents reliably requires robust methods to verify that reported claims are supported by sufficient evidence and that all committed experiments are fully covered by retained records. Traditional logging and hash-linked transcripts often fall short in providing this level of assurance. This paper introduces ClaimReceipt, a novel solution designed to address these evidentiary challenges. ClaimReceipt is a claim-relative receipt specification and a selective verifier. It works by binding specific, typed transaction evidence to a signed experiment manifest, allowing for clear audit verdicts of PASS, INVALID, or INCONCLUSIVE for each claim. The specification was frozen before implementation to ensure integrity. In tests, a ClaimReceipt verifier accurately reproduced manual audit verdicts on historical records, replayed deterministic and post-generation records, and correctly identified semantic faults without false positives. A prospective epoch demonstrated its ability to ensure coverage and accounting, with specific outcomes for withholding terminal receipts or private evidence, matching preregistered predictions. The system adds minimal overhead, highlighting its practical applicability for enhancing the rigor and trustworthiness of AI agent evaluations.

Why it matters

For professionals involved in AI development, auditing, or compliance, ClaimReceipt offers a critical tool for establishing trust and transparency in agent evaluations. It ensures that performance claims are verifiable and that experiments are fully documented, which is essential for regulatory scrutiny and responsible AI deployment.

How to implement this in your domain

  1. 1Integrate ClaimReceipt's specification into AI agent development pipelines to generate verifiable evidence for all agent actions and outcomes.
  2. 2Implement the ClaimReceipt verifier to automatically audit agent evaluation results for evidence sufficiency and experiment coverage.
  3. 3Establish clear protocols for signing experiment manifests and encrypting private evidence for secure and auditable agent deployments.
  4. 4Train development and QA teams on the ClaimReceipt framework to ensure consistent application and interpretation of audit verdicts.

Original post by Peiying Zhu, Sidi Chang

"arXiv:2609.01992v1 Announce Type: new Abstract: Agent evaluations face two distinct evidentiary questions: whether a reported claim is recomputable from retained evidence (sufficiency), and whether the retained records cover the committed experiment set (coverage). Generic logs a…"

View on X

Originally posted by Peiying Zhu, Sidi Chang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses