MMLongBench-Doc-V2 Improves Long Document QA Benchmark Accuracy

Mingtian Zhang· August 5, 2026 View original

Key takeaways

  • Benchmark quality significantly impacts the perceived performance of LLMs.
  • MMLongBench-Doc-V2 offers a more accurate and semantically aware evaluation for long-document QA.
  • Correcting ground-truth annotations is vital for reliable benchmark results.
  • LLM judges can provide more nuanced and semantically robust evaluation than string matching.

Who benefits

AI ResearchSoftware DevelopmentData SciencePublishing

Summary

MMLongBench-Doc-V2 is a revised long-document QA benchmark that corrects 106 ground-truth annotations and replaces string-based answer comparison with an LLM judge for semantic evaluation. These changes address issues in the original benchmark that skewed performance measurements for long-context language models.

The MMLongBench-Doc benchmark, used for evaluating long-document Question Answering systems, had significant flaws that distorted performance metrics. These included numerous incorrect or ambiguous annotations and a rigid string-matching evaluation method. The new MMLongBench-Doc-V2 addresses these issues by meticulously correcting over a hundred annotations, providing clear justifications for each change. Furthermore, it shifts from exact string comparison to a semantic evaluation using a pinned LLM judge, which more accurately assesses whether a model's response conveys the correct meaning.

Why it matters

Accurate benchmarks are crucial for developing and comparing long-context language models; this revision ensures more reliable and meaningful evaluation of their capabilities.

How to implement this in your domain

  1. 1Review current LLM evaluation pipelines to ensure they use up-to-date and semantically robust benchmarks.
  2. 2Integrate MMLongBench-Doc-V2 into evaluation suites for long-document QA models.
  3. 3Adopt LLM-based semantic judges for evaluating model outputs, moving beyond strict string matching.
  4. 4Analyze model performance on the revised benchmark to identify true strengths and weaknesses in long-context understanding.
  5. 5Contribute to community efforts for benchmark improvement by reporting issues and suggesting corrections.

Original post by Mingtian Zhang

"arXiv:2608.03397v1 Announce Type: new Abstract: MMLongBench-Doc is a long-document QA benchmark of 1,082 questions over 135 PDFs. Two properties of it push measured scores away from the quantity they are meant to capture: the reference metric compares extracted answers, so 1,358,…"

View on X

Originally posted by Mingtian Zhang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses