Endless Exam measures AI math progress with scalable benchmarks
The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
Artificial Intelligence
Summary
Measuring how well AI models can solve complex math problems is hard, especially as problems get bigger. This paper introduces the Endless Exam, a set of math challenges that grow in size and difficulty to track AI progress toward very advanced intelligence. The authors provide tools that check answers automatically and score them relative to the best known solutions, helping compare different AI systems fairly. They tested eight AI models on many problems but found none beat the current best human or machine results. Their system lets others keep testing future AI as it improves beyond what humans can do.
What this means in practice
- •For ai developers: Compare and track AI models’ math problem-solving progress using scalable challenge families with automatic verification and scoring.
- •For math software engineers: Incorporate scalable, verifiable benchmarks to test math reasoning components across a wide range of problem sizes automatically.
Authors
Muhan Zhang
Abstract
We introduce the Endless Exam, a benchmark for measuring mathematical progress from today's models toward artificial superintelligence through fourteen parameterised construction families. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at $1$. The families draw on open mathematical problems for long-term targets and generate new instances at larger parameters, where compact certificates keep large constructions verifiable. Across eight models evaluated on 69 distinct instances, continuous quality scores distinguish performance even though no evaluated system surpasses a published frontier. Size-quality curves show how construction quality changes as problem size increases. We release the generators, verifiers, references, model responses and analysis to support continued measurement before and beyond human frontiers.