Two AI Metrics Diverged: Will it Make All the Difference?

2026-07-01Artificial Intelligence

Artificial Intelligence
AI summary

The authors examine whether top AI models will always be better than smaller, cheaper models or if simpler models will catch up over time. They find this depends on how AI performance is measured: some metrics show the gap shrinking, while others show big models maintaining or increasing their lead. They identify that metrics with limits (bounded) often allow smaller models to compete, but those without limits (unbounded) favor large, costly models. The authors argue understanding which metrics apply is important for policy decisions, because it affects who will hold advanced AI capabilities—either a few rich actors or many users with modest resources.

AI capabilitiesperformance metricsbounded metricsunbounded metricsvalidation losscompute scalingfrontier modelsmeek modelspolicy implicationstraining compute
Authors
Alex Fogelson, Zachary A. Brown, Hans Gundlach, Jayson Lynch, Neil Thompson
Abstract
As exponential compute scaling continues, will the capabilities of frontier AI models outstrip what is accessible to developers on a small fixed budget? Or will capabilities converge, with "meek models inheriting the earth"? Building on Gundlach et al. (2025b), we show that the answer depends on how we value and measure AI capabilities. We discuss conventional performance measures and show that, while validation loss shows a shrinking gap, on other metrics frontier models grow their lead forever. Classifying performance metrics by their functional forms in relation to training (and inference) compute, we provide tight mathematical conditions for determining which metrics favor meek models, and show that bounded performance metrics always do. But careful interpretation of performance metrics is essential: we show that many common bounded metrics have closely-related counterpart metrics that are unbounded (and vice versa). Determining the apt metric in a domain is a prerequisite for policy, since bounded and unbounded metrics may suggest opposing policy responses. If a particular capability -- like software engineering, synthetic biology, or rhetorical persuasiveness -- is unbounded when measured in the terms we care about, frontier-level capability will likely be concentrated in the hands of a few wealthy actors. Conversely, if that capability is instead bounded, frontier-level capabilities proliferate through meek models into the hands of the many.