Llm judge saves cost by escalating unsure cases for review

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Artificial Intelligence

Summary

Checking if language models can decide which answers are right or wrong can be expensive and sometimes unsure. The authors studied a simple model that only gives a yes or no answer cheaply and raises tricky questions to a more expensive system. They found this two-step plan is almost as accurate but much cheaper, as it only asks for help when unsure. This approach can make evaluating many answers faster and cost less.

What this means in practice

  • For ai model evaluators: Reduce costs by using a fast decision-only judge that escalates uncertain evaluations to a stronger judge for accuracy.
  • For content moderation teams: Speed up automated review by accepting confident decisions automatically and escalating uncertain cases for human or stronger model review.

Authors

Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman

Abstract

LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generative and reward-model judges, with blinded human adjudication, we find it within three percentage points of a state-of-the-art LLM judge, our strongest comparator, on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee. Larger gaps arise when judgments require checking a derivation or resisting an elaborately written wrong answer. On several benchmarks, JEV's gap to this comparator is concentrated in low-confidence decisions. A frozen cascade that accepts confident verdicts and escalates uncertain ones retains 99% of the comparator's accuracy at lower cost.