Summary
Scoring answers in education often involves multiple raters who might disagree, especially when categories are not simply correct or incorrect. This paper presents a new model that measures how much raters agree on different score categories by using AI-derived language similarities to weight their judgments. The model works well even when scores are unevenly distributed and does not assume equally spaced or ordered scoring levels. The authors show that the model mostly confuses nearby scores, preserving the natural order in scoring guides. They also introduce ways to adjust how language similarities are computed to better fit different types of scoring tasks.
Ising modelPotts modelmultinomial datarater reliabilitylanguage model embeddingspairwise agreementscoring rubricsemantic similarityeducational assessmentpower transformation
Abstract
The Ising model is extended to the Potts model for multinomial data. We introduce a Rater Ising-Potts model that uses agreement indicators between pairs of raters and category labels, with weights derived from LLM embeddings. The model does not presuppose ordered category thresholds or equidistant scoring; instead, it focuses directly on pairwise agreement among raters and assigns category-specific positive weights, making it particularly suited for multi-category scoring reliability when raters evaluate responses using a scoring guide. We demonstrate the model's effectiveness on diverse constructed-response tasks, including balanced short-answer items and more challenging, imbalanced essay prompts from the AERA dataset. Across these settings, the model achieves strong agreement with human scores, with the vast majority of misclassifications occurring between adjacent score levels, confirming its ability to preserve the ordinal structure of scoring rubrics without imposing rigid assumptions. A practical similarity normalization and optional power transformation is introduced as a tunable preprocessing step that sharpens semantic distinctions and can be adapted to different datasets. These findings suggest that LLM-derived semantic similarities, combined with this parsimonious Potts-type formulation and flexible similarity scaling, offer a robust and interpretable framework for reliability auditing in educational assessment contexts. Extensions to multiple raters and hierarchical rating processes are discussed.