DocTrace Enhances Traceable Long Document VQA with Evidence Graphs

Le Xiang, Zhicheng Guan, Hong Chen, Xiaocong Lin, Zhenghua Lei, Teng Hu, Bolei He, Long Zeng· August 5, 2026 View original

Key takeaways

  • LongDocVQA requires explicit mechanisms for evidence representation and verification.
  • DocTrace uses hierarchical evidence graph reasoning for transparent MLLM answers.
  • It significantly improves accuracy on long-document benchmarks.
  • The framework provides explicit node-level provenance, enabling verifiable reasoning.

Who benefits

LegalHealthcareFinanceGovernmentResearch

Summary

DocTrace is a hierarchical framework that improves Long Document Visual Question Answering (LongDocVQA) by explicitly representing and verifying evidence through graph reasoning. It localizes evidence, parses documents, and builds evidence graphs, leading to higher accuracy and transparent, verifiable reasoning for MLLMs.

Multimodal Large Language Models (MLLMs) struggle with Long Document Visual Question Answering (LongDocVQA) because they often lack explicit mechanisms to show how they arrive at an answer from diverse document elements. DocTrace addresses this by reframing LongDocVQA as an explicit evidence graph reasoning problem. This hierarchical framework first localizes relevant evidence across multiple pages, then performs structured document parsing, and finally constructs an evidence graph. This graph explicitly links pieces of evidence, allowing for transparent and verifiable reasoning. Through a two-stage training process, DocTrace significantly outperforms existing baselines and proprietary MLLMs, providing not just better answers but also clear provenance for those answers.

Why it matters

For professionals needing to extract precise information and understand the reasoning behind answers from complex, multi-page documents, DocTrace offers a path to more accurate and auditable AI systems.

How to implement this in your domain

  1. 1Assess current MLLM capabilities for long-document understanding and identify areas lacking traceability.
  2. 2Explore integrating evidence graph reasoning into document processing pipelines for critical applications.
  3. 3Develop methods for explicit evidence localization and structured parsing of heterogeneous document elements.
  4. 4Implement a two-stage training approach (SFT followed by GRPO) for MLLMs to enhance evidence-based reasoning.
  5. 5Design user interfaces that visualize the evidence graphs, allowing human experts to verify and audit AI-generated answers.

Original post by Le Xiang, Zhicheng Guan, Hong Chen, Xiaocong Lin, Zhenghua Lei, Teng Hu, Bolei He, Long Zeng

"arXiv:2608.03292v1 Announce Type: new Abstract: Long Document Visual Question Answering (LongDocVQA) requires Multimodal Large Language Models (MLLMs) to locate, integrate, and reason over heterogeneous document elements distributed across multiple pages. Existing approaches, inc…"

View on X

Originally posted by Le Xiang, Zhicheng Guan, Hong Chen, Xiaocong Lin, Zhenghua Lei, Teng Hu, Bolei He, Long Zeng on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses