LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora

Computation and Language

Summary

The authors studied how large language models (LLMs) grade essays and found that these models act like human raters who can differ in strictness and show bias. They tested LLM grading consistency and found big differences in how harsh the models scored essays, larger than differences between trained humans. They also checked if grading changed over time or between versions and found some shifts but no mysterious silent changes. Overall, the LLMs were consistent but less accurate than humans, and the expected bias called 'halo' was not worse than human graders. This work helps understand the limits and variability of using LLMs for essay scoring.

Authors

Veerendra Kumar Sunkavalli

Abstract

Large language models (LLMs) are increasingly used as essay graders in learning analytics, evaluated almost exclusively with agreement statistics. Educational measurement warns that raters also differ in severity, show halo, and drift as instruments. We treat LLM judges as raters and run a pre-registered rater-effects battery (many-facet Rasch severity, residual halo, generalizability/decision studies, cross-version shifts, differential functioning) on public corpora in two languages (ENEM/Essay-BR; ASAP): 2,377 essays, 12 judges, 4 providers, 5 version contrasts, replicated cells, released as a score tensor. Judge severity spans 219 points on ENEM's 0-1000 scale; on ASAP the panel spread is 15-33% of the score range against a between-trained-human gap near 1%. Judge-human correlations sit in an undiscriminating .47-.56 band. All five version contrasts shift severity beyond a family-wise permutation null (up to 133 points), and one judge was deprecated mid-study, caught by identity canaries. Two pre-registered tests returned honest nulls: severity-adjusted leaderboard reversals did not survive a permutation null, and "silent drift" was refuted: agreement moved with severity in four of five contrasts. Replication yields self-consistency (phi>=.80 at k<=2) but not human-level accuracy, and a same-instrument check overturned our own halo comparison: matched on instrument and calibration, we find no credible evidence that judge halo exceeds the trained-human range.