Papers for

automated grading teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Multimodal verifier improves checking scientific images with explanations

SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation

Abstract: In realistic education, a solution is often expressed not only in words but in a drawing--a circuit, a geometric construction, a function plot--and a teacher must grade the drawing as carefully as the text. Recent advances in unified multimodal models have enabled scientific image generation, yet verifying the correctness of these specialized visual outputs remains a critical bottleneck: errors often arise from intricate domain knowledge, structural reasoning, and multi-step instruction rather than surface-level artifacts. Existing verifiers mainly target natural images and compress judgement into scalar scores, leaving scientific coverage and explainable feedback for error correction underexplored. To bridge this gap, we make three main contributions. (1) We construct SciGen-Verify, a benchmark dedicated to explainable verification of scientific image generation, spanning instruction following, multidisciplinary reasoning, and world knowledge domains. It contains a three-tier hierarchical protocol over the binary judgement, supporting explanation, and corrective editing instruction. (2) We develop SciGen-Verifier, a reasoning-driven multimodal verifier trained via cold-start supervised fine-tuning followed by a curriculum-based two-stage reinforcement learning pipeline. The rubric-guided process rewards first strengthen scientific reasoning exploration and outcome rewards subsequently align output with ground-truth annotation. (3) On SciGen-Verify, SciGen-Verifier achieves competitive performance against much larger proprietary models. It further serves as a practical online critic for iterative image rectification.

Sun 27 SeptComputer Vision and Pattern Recognition
The gist
School assignments often include drawings like circuits or graphs, but machines find it hard to check if these pictures are correct. The authors created a special test and a tool called SciGen-Verifier that can look at scientific images, understand instructions, and explain whether they are right or wrong. Their system learns step by step to reason carefully about the images and gives helpful feedback so mistakes can be fixed. It works well compared to bigger computer models and can help improve images by pointing out errors.
Open → 2609.33399v1

Typed classifier judges answers cheaper and faster than llms with similar mistakes

JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places

Abstract: We ask whether Jev, a typed classifier that returns probabilities over permitted answers without generating text, can replace an LLM rubric judge. We compare it with three flash-tier LLM judges on nine panels drawn from seven benchmarks, giving every judge identical criterion texts. Jev's accuracy differs significantly from an LLM judge's in only 8 of 27 paired comparisons, ahead mostly on binary criteria and behind only on graded ones, and most of the other comparisons are inconclusive. Summed over the nine panels, the LLM judges, called once per criterion, cost 29 to 325 times as much as Jev and took 30 to 220 times as long. On graded criteria all four judges agree more with one another than with the labels and mostly assign lower levels than the raters. One of several observational accounts is that raters followed scale conventions our criterion texts omit. Jev's confidence ranks its own errors on most panels, which should make a cheap classifier the ideal first stage of a cascade that defers its uncertain verdicts to an LLM judge. Correlated errors undo that advantage. The LLM judges repeat nearly all of Jev's most confident errors, so a cascade replayed on the recorded verdicts lowers cost but gains at most 1.5 points over the best single judge with cross-fitted thresholds, and at most 2.0 even with oracle thresholds.

Thu 24 SeptComputation and Language
The gist
The authors investigate whether a typed classifier called Jev can be used instead of large language models (LLMs) to judge answers based on rubrics. Jev is much cheaper and faster than LLM judges and performs similarly in accuracy, especially on simple yes/no criteria. However, both Jev and LLM judges tend to make similar mistakes, especially on more graded, complex criteria. This means using Jev first and then asking an LLM only on uncertain answers saves cost and time but improves accuracy only a little.
Open → 2609.29769v1