DocMemo Enhances Multi-Modal Long Document Understanding

Hanshu Yao, Janfeng Zhong, Niu Lian, Jinpeng Wang· August 10, 2026 View original

Key takeaways

  • DocMemo improves long-document understanding by dynamically discovering evidence.
  • It uses a tri-level memory system to track document structure, page relevance, and query context.
  • Dynamic page belief updating and multi-modal evidence access enhance accuracy.
  • This framework achieves state-of-the-art performance on complex document understanding tasks.

Who benefits

LegalBFSIHealthcareGovernmentResearch

Summary

DocMemo is a memory-guided framework that improves long-document understanding by dynamically discovering evidence across hundreds of pages using probabilistic memory-guided retrieval. It addresses limitations of static retrieval and fragile cross-round memory in existing systems.

Understanding long, multi-modal documents, especially those spanning hundreds of pages, presents a significant challenge for AI systems. Existing methods often struggle with static retrieval, committing to a fixed set of pages early on, or with iterative approaches that lack robust mechanisms for tracking dynamic changes in page relevance. This can lead to errors that are difficult to recover from.Researchers have introduced DocMemo, a novel framework that redefines long-document reasoning as a dynamic exploration of evidence. DocMemo employs a sophisticated tri-level retrieval state, comprising Document Schema Memory, Page Belief Memory, and Question Episodic Memory. These components respectively capture structural priors, dynamically estimate page relevance, and track query-specific reasoning paths.During the reasoning process, DocMemo continuously refines its cross-round page selection. It achieves this through Bayesian page belief updating, Thompson sampling, spatial proximity propagation, and structure-aware adaptive-granularity evidence access. Furthermore, it augments page-level evidence with fine-grained visual regions, crucial for multi-modal understanding. Experiments across three benchmarks demonstrate DocMemo's state-of-the-art performance, validating the effectiveness of its structured memory and dynamic belief updating mechanisms.

Why it matters

This advancement significantly improves AI's ability to accurately process and extract information from complex, lengthy documents, which is critical for many enterprise applications.

How to implement this in your domain

  1. 1Evaluate DocMemo's capabilities for internal document processing needs, especially for legal, financial, or research documents.
  2. 2Integrate dynamic evidence discovery mechanisms into existing information retrieval or knowledge management systems.
  3. 3Develop custom document schemas to leverage DocMemo's Document Schema Memory for specific organizational data.
  4. 4Pilot DocMemo or similar memory-guided retrieval systems for complex query answering over large document repositories.

Original post by Hanshu Yao, Janfeng Zhong, Niu Lian, Jinpeng Wang

"arXiv:2608.07067v1 Announce Type: new Abstract: Long-document understanding requires locating sparse and heterogeneous evidence across hundreds of pages, yet existing systems remain limited by static retrieval and fragile cross-round memory. Mainstream single-round methods commit…"

View on X

Originally posted by Hanshu Yao, Janfeng Zhong, Niu Lian, Jinpeng Wang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses