APEX-Accounting

2026-07-29Computation and Language

Computation and LanguageArtificial IntelligenceHuman-Computer Interaction
AI summary

The authors created APEX-Accounting, a test to see if advanced AI models can handle real accounting tasks like balancing accounts and making reports. Experts designed and checked 160 detailed tasks using various file types to simulate real accounting work. They tested nine top AI models, but none performed perfectly; the best models achieved around 50-56% accuracy on major criteria. They also found a surprising pattern where models seemed to do better with more allowed words overall, but spent more words on harder tasks, resulting in lower scores within the same word limits. The benchmark is private, but the authors offer to evaluate new models upon request.

APEX-Accountingaccounting tasksbenchmarkfrontier modelstoken budgetSimpson's paradoxMean Criteria@3Pass@8Claude-Fable-5Muse-Spark-1.1
Authors
Julien Benchek, Austin Bennett, Jasmin Kern, Ryan Stevens, Rene Sultan, Charis Ching, Hayley Popiel, Vaibhav Mittal, Felix Mercier, Brendan Foody, Bertie Vidgen
Abstract
We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, ahead of Muse-Spark-1.1 (xHigh) at 52.6%. No model scores more than 2.6% Pass^8 (GPT-5.6-Sol (Max+Pro)) and the highest Pass@8 is 21.5% (Muse-Spark-1.1 (xHigh)). We experiment with increasing the token budget from $1 to $50 and observe an instance of Simpson's paradox: scores increase as the token budget increases but within a given budget-constrained harness, scores are lower on tasks where the model spends more tokens. As APEX-Accounting is a closed benchmark, leaderboard evals can be run for any frontier model on request.