Multimodal verifier improves checking scientific images with explanations

SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation

Computer Vision and Pattern Recognition

Summary

School assignments often include drawings like circuits or graphs, but machines find it hard to check if these pictures are correct. The authors created a special test and a tool called SciGen-Verifier that can look at scientific images, understand instructions, and explain whether they are right or wrong. Their system learns step by step to reason carefully about the images and gives helpful feedback so mistakes can be fixed. It works well compared to bigger computer models and can help improve images by pointing out errors.

What this means in practice

Authors

Jiali Chen, Zhengteng Lin, Zuqi Wang, Shirong Lin, Xi Yu, Xusen Hei, DingBa Fu, Jiayuan Xie, Yi Cai

Abstract

In realistic education, a solution is often expressed not only in words but in a drawing--a circuit, a geometric construction, a function plot--and a teacher must grade the drawing as carefully as the text. Recent advances in unified multimodal models have enabled scientific image generation, yet verifying the correctness of these specialized visual outputs remains a critical bottleneck: errors often arise from intricate domain knowledge, structural reasoning, and multi-step instruction rather than surface-level artifacts. Existing verifiers mainly target natural images and compress judgement into scalar scores, leaving scientific coverage and explainable feedback for error correction underexplored. To bridge this gap, we make three main contributions. (1) We construct SciGen-Verify, a benchmark dedicated to explainable verification of scientific image generation, spanning instruction following, multidisciplinary reasoning, and world knowledge domains. It contains a three-tier hierarchical protocol over the binary judgement, supporting explanation, and corrective editing instruction. (2) We develop SciGen-Verifier, a reasoning-driven multimodal verifier trained via cold-start supervised fine-tuning followed by a curriculum-based two-stage reinforcement learning pipeline. The rubric-guided process rewards first strengthen scientific reasoning exploration and outcome rewards subsequently align output with ground-truth annotation. (3) On SciGen-Verify, SciGen-Verifier achieves competitive performance against much larger proprietary models. It further serves as a practical online critic for iterative image rectification.