CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

2026-08-27Artificial Intelligence

Artificial IntelligenceComputation and LanguageInformation RetrievalMachine Learning
AI summary

The authors created CorporateBench (CB), a new test to see how well large language models (LLMs) can answer questions about really big sets of company documents. Unlike past tests, CB uses huge, realistic collections from fake companies, ensuring that the information is consistent over time and across many documents. They tested five different LLMs and found that as the amount of information grows to real company sizes, the models struggle more. CB helps researchers better measure how well these models understand and work with complex corporate information.

Large Language ModelsEnterprise Document CollectionsQuestion Answering BenchmarkInformation ExtractionKnowledge Base QueryingSynthetic DataCross-document ConsistencyCorporate CommunicationModel Evaluation
Authors
Sil Hamilton, Albert Yu Sun, Oscar J. Romero, Carl-Leander Henneking, David Mimno, Bishan Yang, Igor Labutov
Abstract
LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. We present CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approaches the conditions LLMs encounter in corporate communication networks, with evaluation corpora surpassing 230,000 documents. CB evaluates LLMs across two dimensions (information extraction and knowledge base querying) through four synthetically generated firms ranging from 12 to 10,000 employees. Each corpus is sampled from a temporally evolving knowledge base describing a consistent world, guaranteeing cross-document logical consistency even across hundreds of thousands of documents. We evaluate five LLMs on CB, revealing increasingly poor performance as input size approaches realistic scales. CB provides LLM developers a metric for corporate communication reasoning, filling a crucial gap in the benchmarking ecosystem.