Papers for

ai model evaluators

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Function level scores measure flag rate more than vulnerability detection quality

A Function-Level Vulnerability Score Measures Flag Rate More Than the Model: Protocol Effects on Paired Benchmarks

Abstract: Language models are increasingly evaluated as vulnerability detectors, and scores reported for similar models differ widely between papers. We measured how much of that difference evaluation protocol accounts for, with model outputs held fixed. In a paired test, a model must flag a vulnerable function and clear its version after a fixing commit. Three choices that published evaluations make differently were varied one at a time: metric, verdict extraction and output budget. Seven frontier and large open models were evaluated on five released pair benchmarks and a set pooled for this work under one protocol, and 61 open models of 1.5B to 36B parameters on the pooled set. Function-level F1 follows how often a model flags both functions of a pair (Spearman $+0.86$ over 42 combinations) and is nearly unrelated to pair-level correctness ($+0.16$). On the pair score, extraction changes a model's number by $+0.001$ at the median and budget by $+0.02$ with an interval through zero, whereas the model changes a benchmark's number by up to 0.18 and the benchmark a model's by up to 0.16; a function-level score therefore measures flag rate more than model. For 37 of 68 models the difference between correct and reversed pairs is within its 95% interval of zero, the value for a null model that flags each function at a fixed rate, while both-flagged and both-cleared rates exceed that null by 0.055 on median, and for 64 of 68 both functions of a pair receive one answer more often than independence predicts: verdicts are determined by the text common to both functions. On length-matched pairs a linear probe on activations separates 0.78 by within-pair ranking, against 0.64 for a tf-idf baseline and 0.5 for length; the generated verdict is near chance for three of six models and at 0.55 to 0.57 for the other three, and a prompted logit is at chance for all six.

Sat 26 SeptMachine Learning
The gist
Detecting software vulnerabilities using AI language models involves checking if the model flags functions as risky and then clears them after fixes. The authors found that the way these models are evaluated affects the reported scores more than actual model differences. They show many scores mostly reflect how often the model flags code rather than how accurately it detects vulnerabilities. Also, judgments about paired code snippets tend to depend on shared text rather than deeper understanding.
Open → 2609.32890v1

Emoji rating bias hides true differences among language models

Two Emojis of Difference: What Multilingual Affective Generation Benchmarks Actually Measure

Abstract: We audit a multilingual affective generation benchmark eight instruction-tuned LLMs producing emoji summaries for 17,100 Bangla, English and Hindi sentences, with 6,960 human judgements and find its headline conclusions to be artefacts of the measurement instrument rather than properties of the systems. Treating annotators as a random rather than a fixed factor, no system differs significantly from any other ($F(7,14)=0.59$, $p=0.76$), although the conventional analysis declares 19 of 28 pairwise differences significant. Annotator identity explains far more rating variance than system identity, and the winning system changes whenever any single annotator is removed. The ordering that does emerge tracks output length: mean emoji count explains 78.7\% of between-system variance, and a within-item length-matched comparison over 2,599 pairs reverses the leaderboard. We further show that cross-provider anisotropy differences vanish under mean-centring, that per-language token costs change sign with the normalising unit, and that multi-view row-wise splits inflate macro-F1 by $3.1$ points and change the top-ranked system. In place of preference scoring we propose **emoji-affect decodability**, a reference-based probe whose rankings are stable to $\pm0.003$ macro-F1 across seeds.

Thu 24 SeptComputation and Language
The gist
Measuring how well language models generate emotional emoji summaries in multiple languages is tricky because human ratings vary a lot depending on the reviewer. The authors show that differences between language models mostly disappear when factoring in individual annotator preferences. They also find the number of emojis used affects scores more than quality. Instead of current rating methods, the authors propose a more stable way to measure emoji-based emotion generation.
Open → 2609.29445v1

Llm judge saves cost by escalating unsure cases for review

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Abstract: LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generative and reward-model judges, with blinded human adjudication, we find it within three percentage points of a state-of-the-art LLM judge, our strongest comparator, on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee. Larger gaps arise when judgments require checking a derivation or resisting an elaborately written wrong answer. On several benchmarks, JEV's gap to this comparator is concentrated in low-confidence decisions. A frozen cascade that accepts confident verdicts and escalates uncertain ones retains 99% of the comparator's accuracy at lower cost.

Tue 22 SeptArtificial Intelligence
The gist
Checking if language models can decide which answers are right or wrong can be expensive and sometimes unsure. The authors studied a simple model that only gives a yes or no answer cheaply and raises tricky questions to a more expensive system. They found this two-step plan is almost as accurate but much cheaper, as it only asks for help when unsure. This approach can make evaluating many answers faster and cost less.
Open → 2609.26550v1

De-identified résumés still reveal ethnicity beyond language details

Beyond the Name: Demographic Leakage in De-Identified Résumés and Evaluation Artifacts in LLM Bias Audits

Abstract: De-identified résumé screening assumes that redacting explicit fields prevents ethnocultural inference; however, recent audits attribute residual leakage to declared languages. We investigate whether eliminating language fields resolves this leakage across nine open-weight models and 620 counterfactual résumés. By holding language attributes strictly identical, we isolate unstructured prose across five ethnocultural conditions and three cue-salience tiers. Target-group recovery averages 0.757 overall and saturates at 1.000 under high salience, demonstrating that non-language prose sustains demographic inference. Crucially, models diverge only under faint cues (0.086-0.690), establishing salience as an essential evaluation axis. Furthermore, pairwise LLM-as-a-judge outcomes are highly sensitive to evaluation design: forbidding ties yields an apparent selection-rate ratio of 0.39 alongside strong position and content effects, whereas permitting ties produces near-universal ties for most models ($\ge94\%$). Downstream scoring shows only very small between-condition differences, highlighting the need to distinguish demographic signals recoverable from résumé content from effects introduced by the evaluation protocol.

Tue 15 SeptComputation and LanguageArtificial Intelligence
The gist
Removing obvious personal information like language skills from résumés does not fully stop AI models from guessing a person’s ethnic background. The authors tested multiple AI models on résumés where they controlled language details to be the same but changed other writing aspects. Even without language clues, the AI could often identify ethnic groups, especially when subtle hints in the text were strong. They also found that how evaluation tests are designed can greatly change the results, showing that measuring bias needs careful methods.
Open → 2609.16501v1