New Protocol Guards LLM Citation Faithfulness in Science.

Taewan Goo, Junsik Kim, Kyulhee Han, GwonYul Jo, Jong-Soo Kim, Tae-Hyung Kim· July 24, 2026 View original

Summary

This research introduces a gold-anchored evaluation protocol and a deployable guard to measure and bound unsupported citations in agentic LLM scientific synthesis systems. It reveals that citation verification is unreliable and inconsistent across verifiers, providing a robust method to ensure citation faithfulness.

New research addresses the critical issue of citation faithfulness in agentic Large Language Model (LLM) systems designed for scientific synthesis, such as OpenScholar and PaperQA2. The study highlights that existing methods for checking citation validity are unreliable, with measured unsupported-citation rates varying significantly (3% to 18%) depending solely on the verifier's strictness. Furthermore, while verifiers might agree on supported citations, they disagree on which ones to flag as unsupported, making cross-paper comparisons invalid without a standardized protocol. To counter this, the researchers propose a gold-anchored evaluation protocol and a deployable guard. The protocol validates the verifier itself, measures re-attribution, and calibrates guarantees against human-annotated gold standards rather than other model verdicts. It allows for swappable verifiers chosen based on cost and demonstrates that a deterministic BM25 approach can match the best open generators for re-attribution. The guard adds a split-conformal layer, providing a distribution-free, finite-sample bound on truly unsupported citations that bypass flagging rules. This offers a quantifiable guarantee on the catch rate. The protocol and guard have been validated across four open 27-35B models and three agentic pipelines on public benchmarks, shipping as an open single-GPU kit.

Why it matters

For professionals relying on AI for scientific literature review, research synthesis, or knowledge extraction, this work provides essential tools to ensure the trustworthiness and factual accuracy of AI-generated content, mitigating the risk of propagating misinformation.

How to implement this in your domain

  1. 1Adopt the proposed gold-anchored evaluation protocol to rigorously assess the citation faithfulness of your LLM-based scientific synthesis tools.
  2. 2Integrate the deployable guard into your agentic LLM pipelines to provide quantifiable bounds on unsupported citations.
  3. 3Standardize your citation verification processes by selecting and calibrating verifiers based on cost and performance against human gold standards.
  4. 4Educate your team on the limitations of current LLM citation verification and the importance of robust validation frameworks.

Who benefits

Research & AcademiaPharmaceuticalsLegalHealthcareAI/ML Development

Key takeaways

  • Existing LLM citation verification methods are unreliable and inconsistent.
  • A new gold-anchored protocol and deployable guard enhance citation faithfulness measurement.
  • The guard provides a quantifiable, distribution-free bound on unsupported citations.
  • Implementing these tools is crucial for trustworthy AI-driven scientific synthesis.

Original post by Taewan Goo, Junsik Kim, Kyulhee Han, GwonYul Jo, Jong-Soo Kim, Tae-Hyung Kim

"arXiv:2607.20527v1 Announce Type: new Abstract: Agentic LLM systems such as OpenScholar and PaperQA2 read the scientific literature and return cited answers, and both they and their benchmarks already check whether those citations hold, with a fixed attribution model or human gra…"

View on X

Originally posted by Taewan Goo, Junsik Kim, Kyulhee Han, GwonYul Jo, Jong-Soo Kim, Tae-Hyung Kim on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses