Visual grounding improves with confidence awareness to reduce hallucinations
SAVOR: Self-Aware Visual Grounding via Confidence-Calibrated Reinforcement Learning for Multimodal Hallucination Mitigation
Computer Vision and Pattern Recognition
Summary
Multimodal large language models sometimes make confident but incorrect claims about images, describing things that aren’t really there. The authors found that teaching these models to be aware of their own confidence helps them avoid making false statements. They created a training method called Savor that lets the model judge when it is uncertain and look again at the image before answering. This approach reduces mistakes while keeping the model’s overall ability to understand images and text.
What this means in practice
- •For ai application developers: Build more reliable image understanding tools that detect when answers may be unreliable and reduce false visual claims.
- •For multimodal system integrators: Enhance multimodal chatbot responses by integrating self-assessed confidence to trigger additional image analysis only if uncertain.
Authors
Zixiu Ding, Zilin Zhao, Yingjie He, Xinlang Kang, Guansu Wang, Wei Zhang
Abstract
Multimodal large language models (MLLMs) have made strong progress on visual question answering and image captioning, yet they still produce fluent claims about objects, attributes, or relations that are not grounded in the image. Many remedies either modify decoding at test time, which adds latency, or fine tune with preferences such as DPO variants, which teach which answer is preferred but not when the model's own answer is unreliable. We argue that calibrated self assessment is the missing signal. We introduce Savor, a training framework that (i) augments the output schema with token and answer confidence, (ii) optimises the policy with a Group Relative Policy Optimisation (GRPO) objective that penalises calibration error and poor abstention decisions, and (iii) uses the learned confidence at inference time to revisit visual evidence only when the model is uncertain. Experiments on POPE, HallusionBench, AMBER and MMHal-Bench across two recent backbones (InternVL3-8B and Qwen3-VL-8B) show that Savor reduces hallucination while preserving general capability on MME and MMBench, with lower Expected Calibration Error than DPO and decoding baselines.