Summary
When groups of judges score something like AI outputs, a separate reference or "anchor" is often used to tell apart true quality from shared mistakes the judges might make together. This paper studies how to estimate if and how these anchors are themselves biased or share errors with the judges, instead of assuming they are perfect. The authors develop a method that works under a simple model involving one main common error factor, requires at least two judges and anchors, and provides mathematical formulas to identify error and quality. They also create tests and tools to check when their method can be trusted and understand possible errors, and show how it can be adapted for different types of scores. Their approach is carefully tested in simulations and applied to real data, where the tests often find the model doesn’t fit perfectly.
AnchorJudge panelError correlationCommon-factor modelQuality signalContaminationDiagnostic batteryOver-identificationBootstrap confidence intervalsOrdinal scores
Abstract
When an external reference set (an anchor) is used to decompose an LLM-judge panel's error into a quality signal and a shared common-mode error, standard practice assumes the anchor is uncontaminated: its error uncorrelated with the judges' shared error. We study when that assumption can be dropped and replaced by an estimate. Under a single-common-factor model, >=2 judges and >=2 anchors point-identify the quality variance, the common-mode variance, and each anchor's contamination correlation rho_k in closed form, with an exact per-anchor-pair failure boundary; a designated clean-anchor estimator, by contrast, reports a contaminated companion anchor as fully clean once its trusted anchor is itself contaminated. Because the single-common-factor assumption is itself untestable, the estimator ships gated behind a calibrated diagnostic battery (judge-covariance dispersion; over-identification; a family-block test from judge metadata, with a family-blocked estimator that removes family-level shared-residual bias exactly), bootstrap confidence intervals with measured coverage, and a weak-identification screen. A proposition maps which violations bias rho_k, in which direction, and which evade detection. For ordinal scores we show an identification hierarchy: with all variables ordinal, rho_k is not identified at any number of anchors; with ordinal judges and >=3 continuous anchors it is, and we give an estimator for that case. On real data the validation is asymmetric, and we say so plainly: the diagnostics are validated in the rejecting direction (both real panels we test are correctly rejected by the model-adequacy pre-test), while the estimator is validated in simulation and stress-tested semi-synthetically under oracle calibration; no real panel has yet passed the pre-test, and the pre-test exists precisely to say so. All results replay offline from shipped, checksummed artifacts.