Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias

2026-08-03Sound

SoundHuman-Computer Interaction
AI summary

The authors studied how well a version of GPT-4o-mini can score detailed student essays about music analysis compared to human teachers. They tested three different ways of prompting the model and found that one method (few-shot prompting with chain-of-thought) matched teacher scores best. Another method tended to give higher scores than teachers, while the third was very consistent but less accurate at matching teacher judgments. The model’s ability to score varied depending on the essay aspect, with terminology being harder to score than reasoning. Overall, the authors suggest that using GPT-4o-mini for scoring is promising but still needs careful setup and human checking.

GPT-4o-minifew-shot promptingchain-of-thought reasoningretrieval-augmented generationself-consistencyrubric-based scoringmusic analysisinter-rater reliabilityKrippendorff's alphaquadratic weighted kappa
Authors
Baicheng Lin, Lingxi Jin, Kyung-Seok Min
Abstract
Scoring open-ended music analysis responses is time-consuming and requires nuanced judgments of harmonic knowledge and formal understanding. This study evaluates the validity and repeatability of GPT-4o-mini for rubric-based scoring of music analysis essays, using teacher mean scores as the benchmark. A dataset of 300 university-level student responses was scored by teachers on four dimensions: Harmony, Form, Reasoning, and Terminology. GPT-4o-mini scored the same responses using three prompting strategies: few-shot prompting with chain-of-thought reasoning (Fs+CoT), retrieval-augmented generation (RAG), and self-consistency based on five internal generations per administration (SC). Each strategy was administered three times with the model, prompt, rubric, and response held constant. Single-pass scores represented an operational scoring condition, whereas median aggregation across three runs was used to examine robustness. Agreement with teacher mean scores was evaluated using correlation, intraclass correlation, Krippendorff's alpha, quadratic weighted kappa, and scoring error indices. Fs+CoT showed the strongest agreement with teacher mean scores in both single-pass scoring and median aggregation. RAG showed systematic over-scoring, whereas SC produced highly repeatable scores but weaker individual-level agreement. Dimension-level analyses showed that scoring performance varied across rubric components, with Terminology generally showing weaker agreement than Reasoning. These findings indicate that GPT-4o-mini can generate stable scores for complex music analysis responses, but prompting strategies produce distinct scoring profiles. Operational use therefore requires strategy-specific calibration, dimension-level validation, and continued human oversight.