Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems
2026-08-19 • Artificial Intelligence
Artificial IntelligenceComputers and SocietyMachine LearningSoftware Engineering
AI summaryⓘ
The authors argue that comparing language models by how often they get the right answer (capability) misses an important point: how consistently they give similar answers each time (precision). They propose measuring this precision by running the same tests multiple times and seeing how much the results vary, without needing complex grading. This measurement helps distinguish between errors that can be fixed easily by adjusting usage from those needing changes in the model itself. Their approach shows that improvement should be judged based on real-world performance, not just constructed test rules.
language modelscapabilityprecisionbenchmarkingconsistencycentral tendencyvariancesampling temperaturemodel evaluationreliability
Authors
George Andrikopoulos
Abstract
Frontier language models are compared, marketed, and benchmarked on capability -- what their best or average output can achieve. I argue this measures the wrong axis. The models have saturated accuracy: their mean output lands on the target. What now separates one system from another in practice is precision: how tightly concentrated their outputs are around that target across repeated, identical requests. Borrowing the marksman's distinction, capability is where the average shot lands; reliability is the size of the group. I make three claims. First, precision, not capability, is the frontier differentiator between systems, and benchmark culture systematically fails to measure it, reporting central tendency rather than spread. Second, precision is measurable, cheaply and without circularity, by running a fixed suite of deterministically scored tasks many times at fixed temperature and computing the per-task consistency of outcomes -- no model-in-the-loop grader required. Third, the measurement is not merely descriptive but decision-guiding: it separates consistent failures (a tight group off-centre, correctable by the operating discipline of Paper 1 -- a sight adjustment) from scattered failures (a wide group, correctable only by changing the model or its sampling -- a rifle problem). I define a grouping metric, specify a harness, and show how tracking a human-AI pair's grouping over time yields the compounding signal that Paper 1's field study requires. A first real run, since replicated, illustrates both the method and its most important limit: one measured gap was closed completely by a single rule (0/5 -> 5/5), while a suite of tasks authored from the rules themselves found no value, because a frontier model already embodies explicit good practice -- establishing that a discipline's worth is found by measurement on real work, not constructed from its own rulebook.