DocHop Benchmarks Multi-Hop Reasoning in Documents

Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park, Xinyi Gu, Zexue He, Soochahn Lee, Rogerio Feris, Yong Jae Lee· September 3, 2026 View original

Key takeaways

  • DocHop benchmarks MLLMs on integrated text and chart reasoning in documents.
  • It highlights a significant performance gap between MLLMs and humans in multi-hop reasoning.
  • Current MLLMs struggle with increasing reasoning complexity in cross-modal tasks.
  • The benchmark provides a controlled testbed for advancing MLLM capabilities.

Who benefits

Financial ServicesLegalHealthcareBusiness IntelligencePublishing

Summary

DocHop is a new benchmark designed to evaluate Multimodal Large Language Models (MLLMs) on integrated chart-context reasoning within information-dense documents. It requires models to perform multi-step compositional reasoning by combining textual narrative with data from multiple charts.

Multimodal Large Language Models (MLLMs) have demonstrated strong capabilities in understanding structured visual information, particularly in tasks like chart and document question answering. However, existing evaluation benchmarks typically assess these domains in isolation, overlooking a crucial ability: how well models can leverage textual context to guide the selection, interpretation, and aggregation of data presented in charts. To address this gap, researchers have introduced DocHop, a novel benchmark specifically created for integrated chart-context reasoning within document-style images. In DocHop, a document's narrative provides multi-step compositional constraints, while embedded charts supply the necessary data values. Questions are grounded in semantic reference labels defined within the narrative, compelling models to first resolve target entities from the context before aggregating evidence across multiple charts. DocHop was constructed using a stochastic logic-first generation pipeline, allowing for controlled reasoning depth and visual density. It comprises 2,074 examples across six task categories. Initial experiments with various proprietary and open-source MLLMs reveal a substantial performance gap compared to human annotators, who achieve over 90% accuracy, while the best model only reaches 62.83%. Models enhanced with reasoning capabilities show improvements, but their performance degrades as reasoning complexity increases, indicating a significant challenge for current MLLMs in complex multi-hop document reasoning.

Why it matters

Professionals developing or deploying MLLMs need to understand their limitations in complex document understanding, especially when integrating text and visual data for multi-step reasoning.

How to implement this in your domain

  1. 1Evaluate current MLLM solutions against the DocHop benchmark to identify reasoning gaps.
  2. 2Prioritize research and development into MLLM architectures that excel at multi-hop, cross-modal reasoning.
  3. 3Develop internal testing methodologies that mimic DocHop's integrated chart-context reasoning challenges.
  4. 4Consider fine-tuning MLLMs on datasets that emphasize complex document understanding tasks.

Original post by Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park, Xinyi Gu, Zexue He, Soochahn Lee, Rogerio Feris, Yong Jae Lee

"arXiv:2609.02059v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isola…"

View on X

Originally posted by Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park, Xinyi Gu, Zexue He, Soochahn Lee, Rogerio Feris, Yong Jae Lee on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses