ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

2026-07-31Artificial Intelligence

Artificial Intelligence
AI summary

The authors created ExtractBench, a new test to check how well computer agents can pull data from different types of business documents using a set format (schema). Their test measures how accurate the extracted information is, how complete the records are, and how well the source of the data can be traced, all while considering the cost of running these agents. They tested various systems and found some perform well on short documents but struggle with longer ones, while others keep accuracy but cost more. A new agent called LlamaExtract Agentic Plus scored best overall, balancing accuracy and cost effectively. The dataset and tools for evaluation are publicly available.

schema-guided extractionenterprise workflowsvalue accuracyrecord completenessgroundingcommercial VLMsorder-insensitive value F1ground-truth curationagentic extractionsource traceability
Authors
Boyang Zhang, Adrian Lyjak, Eli Stewart, Zhaoqi Li, Simon Suo
Abstract
Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We present ExtractBench, a benchmark for schema-guided extraction and, to our knowledge, the first to score value accuracy, record completeness at scale, grounding, and measured cost together. The evaluation system contains 4,869 pages across 370 enterprise documents, 8 business domains, and 67 document types, with clear tags differentiating their challenge scenarios. The scalable schema and ground-truth curation pipeline combines independent-system agreement for real documents, known values for synthetic lists, and human verification for forms. We report order-insensitive value F1 for value accuracy, plus two grounding metrics for source traceability: word- and page-level F1. Commercial VLMs perform well on short documents but often truncate record lists on long ones, while coding agents retain higher accuracy at much higher cost. LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost. Dataset and evaluation code are available on \href{https://huggingface.co/datasets/llamaindex/ExtractBench}{HuggingFace} and \href{https://github.com/run-llama/ExtractBench}{GitHub}.