Right Answer, Wrong Heat: Explanation-Aware Evaluation and Thermal-Grounded Feedback for MLLMs on Infrared Images

2026-08-10Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors study how multimodal large language models (MLLMs) answer questions about infrared images. They find that even if the answer is correct, the model's explanation might not actually use the thermal information from the infrared image. They introduce a method to check if explanations truly rely on thermal data and show that some models use visible-light clues instead. The authors also propose Thermal-Grounded Feedback (TGF), a way to improve explanations without changing the answers. They suggest future models should focus on giving thermally grounded explanations, not just correct answers.

Multimodal Large Language ModelsInfrared ImagingVisual Question AnsweringThermal GroundingExplanation EvaluationDual-LLM Consensus JudgeThermal-Grounded FeedbackVisible Light vs. InfraredModel ExplanationCalibration
Authors
Yongsong Huang, Xiaofeng Liu, Tomo Miyazaki, Yaohou Fan, Shinichiro Omachi
Abstract
General-purpose multimodal large language models (MLLMs) are increasingly applied to infrared images, where they are commonly scored by answer accuracy alone. However, a correct answer does not ensure that the model's explanation is grounded in infrared thermal evidence. We introduce an explanation-aware evaluation framework that separates answer correctness, output-level explanation groundedness, and thermal grounding for infrared visual questions. Using a Dual-LLM Consensus Judge with a preliminary human-anchor calibration check, we find that correct answers can still rely on weak or visible-light evidence; withholding the original infrared image and showing only a visible-like rendering erodes thermal grounding with little accuracy change; and this erosion is observed most strongly for more capable models but disappears when infrared remains available. We further propose Thermal-Grounded Feedback (TGF), a training-free feedback loop that diagnoses explanation-side failures and revises the explanation while preserving the selected answer. On local paired-input validation, TGF improves explanation-side grounding without changing answers. These findings suggest that future trustworthy MLLMs for infrared scene understanding should be evaluated and developed to produce thermally grounded explanations rather than merely accurate answers.