Reviewers stop rewarding complex writing as AI makes it easy
Do Reviewers Still Reward Lexical Complexity? A Frozen-Rater Study of Preference Drift in 124K ICLR Reviews
Computation and LanguageDigital LibrariesMachine Learning
Summary
This paper looks at whether people who review scientific papers still give better scores for complicated writing now that AI can write complex text easily. The authors used a clever method by having an AI model review papers from past years all at once, so changes in scores can only be due to human reviewers' changing preferences. They found that over time, human reviewers stopped giving extra credit for fancy word use, but still rewarded variety in sentence length. The AI model kept valuing complexity like it did in earlier years, showing that human preferences have shifted while AI judgments stayed constant.
peer reviewlexical complexitylarge language modelsICLRfrozen raterpreference driftregression analysisfalse-discovery controlnatural language processingmachine evaluation
Authors
Jiabin Zheng
Abstract
Large language models have collapsed the cost of producing lexically elaborate prose, and whether peer reviewers still reward it is a question about the evaluator, not about the text. When the association between a writing cue and review scores moves across years, the reviewers may have changed, the submissions may have changed, or both, and a regression of scores on text cannot say which. We separate the two with a frozen rater: 81,850 machine reviews of ICLR submissions from 2018 to 2025, all generated in one February-April 2025 window with one model family and one prompt, so that its year-to-year coefficients track submission composition alone and the human-minus-frozen trend difference identifies reviewer preference drift. On 32,638 submissions with 124,615 human reviews, the human coefficient on non-domain lexical complexity falls from +0.142 to -0.015 while the frozen rater moves from +0.080 to +0.082; the three-way difference-in-differences is -0.0100 (q=0.013), and forty random-wordlist placebos through the same specification centre on zero. Humans still reward sentence-length variability, which the frozen rater never registers, while the frozen rater still pays for lexical complexity at its earlier rate. Every claim is held to a double gate of false-discovery control and interval exclusion, and the findings that failed adversarial re-testing are reported. Reviewers discounted a cue whose production cost collapsed, as models of manipulable signals prescribe; an LLM judge calibrated to historical human preferences inherits the earlier schedule and drifts out of alignment while its agreement with humans on totals stays ordinary.