ExtractBench: New Benchmark for Enterprise Document Extraction

Boyang Zhang, Adrian Lyjak, Eli Stewart, Zhaoqi Li, Simon Suo· August 3, 2026 View original

Key takeaways

  • ExtractBench evaluates schema-guided enterprise document extraction.
  • It measures value accuracy, record completeness, grounding, and cost.
  • Commercial VLMs struggle with long documents, while coding agents are costly.
  • LlamaExtract Agentic Plus shows high accuracy at lower cost.

Who benefits

BFSILegalHealthcareGovernmentConsulting

Summary

ExtractBench is introduced as a new benchmark for schema-guided enterprise document extraction, evaluating value accuracy, record completeness, grounding, and cost across 4,869 pages and 370 documents. It reveals that commercial VLMs struggle with long documents, while coding agents are accurate but costly, with LlamaExtract Agentic Plus ranking highest.

This paper presents ExtractBench, a novel benchmark specifically designed to evaluate schema-guided enterprise document extraction. This task is crucial for modern workflows where AI agents must accurately extract information from documents based on a user-defined schema, providing source evidence for grounding. The benchmark is comprehensive, featuring 4,869 pages from 370 enterprise documents across 8 business domains and 67 document types, with clear tags for different challenge scenarios. ExtractBench's evaluation system measures several key aspects: value accuracy (using order-insensitive value F1), record completeness at scale, grounding (with word- and page-level F1 for traceability), and the associated computational cost. Initial evaluations reveal that commercial Vision-Language Models (VLMs) perform well on shorter documents but often truncate record lists in longer ones. Coding agents achieve higher accuracy but at a significantly greater cost. Notably, LlamaExtract Agentic Plus emerged as the top performer across all three metrics, offering accuracy comparable to coding agents at a fraction of the expense.

Why it matters

This benchmark provides a standardized and rigorous way to evaluate the performance of AI agents in a critical enterprise task: extracting structured information from unstructured documents. Professionals can use this to select and optimize AI solutions for document processing, improving efficiency and data accuracy in their operations.

How to implement this in your domain

  1. 1Utilize ExtractBench as a reference to evaluate the performance of existing or new document extraction solutions within your organization.
  2. 2Prioritize AI solutions that demonstrate strong performance on ExtractBench's metrics, especially for handling long documents and ensuring grounding.
  3. 3Investigate LlamaExtract Agentic Plus or similar top-performing agents for enterprise document processing needs.
  4. 4Develop internal testing protocols for document extraction that incorporate metrics like value accuracy, record completeness, and grounding, inspired by ExtractBench.

Original post by Boyang Zhang, Adrian Lyjak, Eli Stewart, Zhaoqi Li, Simon Suo

"arXiv:2607.29677v1 Announce Type: new Abstract: Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as groundin…"

View on X

Originally posted by Boyang Zhang, Adrian Lyjak, Eli Stewart, Zhaoqi Li, Simon Suo on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses