Visual question answering requires knowing when to answer or abstain

When Does an Image Determine the Answer? Benchmarking Visual Answerability across Charts and Scenes

Computer Vision and Pattern Recognition

Summary

Sometimes images don’t have enough clues to answer questions confidently. The authors studied when AI systems should give an answer and when they should say “I don’t know.” They tested this on charts, 3D scenes, and photographs using a new benchmark. Even the best systems often give wrong answers or fail to abstain when needed. Their work shows how evaluating both correct answers and correct refusals reveals unseen problems in these systems.

What this means in practice

  • For ai system developers: Improve visual question answering systems by integrating answer abstention when image evidence is insufficient.
  • For data annotation teams: Use the benchmark to validate dataset quality in visual QA by distinguishing supported answers from guesses.

Authors

Sungguk Cha, Mintae Kim, Youngsub Han, Byoung-Ki Jeon, Sangyeob Lee

Abstract

Reliable visual question answering requires correct answers when evidence is sufficient and abstention when it is not. We introduce a benchmark that connects complete-question evaluation with explicit evidence for its labels across PlotQA charts, CLEVR rendered scenes, and GQA photographs. Each question groups original and edited images, presented independently; success requires every supported answer and every required abstention to be correct. For chart missing-information labels, executable witnesses establish that admissible complete charts give different answers but identical pixels after masking. Scene labels follow source programs and edits, with a residual-cue analysis for photographs. Across 72,000 responses from six model configurations, the highest observed complete task success rates are 57.0%, 43.5%, and 33.7%, respectively. On charts, the strongest configuration achieves 96.2% per-view decision accuracy, yet 265 of its 835 groups with every decision correct still contain incorrect answers. Evaluating supported answers and necessary abstentions together exposes failures that answerability decisions alone conceal.