Fine-Grained Multi Image Object Hallucination Benchmark

2026-08-31Computer Vision and Pattern Recognition

Computer Vision and Pattern RecognitionArtificial IntelligenceMachine Learning
AI summary

The authors created a new test called MIOH to check how well AI models that understand images and text avoid making up things (hallucinations) when looking at multiple pictures. Unlike older tests that focus on single images or general reviews, MIOH carefully checks if models can correctly identify objects, count them, notice their traits, and find their locations across several pictures. They tested 29 AI models and found even the best ones make different kinds of mistakes, especially when combining information from multiple images. The authors show that these mistakes are not just about seeing problems but also about how the models keep track of objects across pictures. MIOH can help improve future AI by providing a clear way to spot and fix these errors.

Multimodal Large Language ModelsObject HallucinationMulti-image ReasoningBenchmarkExistence TaskCounting TaskAttribute TaskPosition TaskVisual ContextPerceptual Difficulty
Authors
Joonki Min, Chaeyun Kim, Hyungwook Choi, Yejin Kim, Kihyun Kim, Yohan Jo, Joonseok Lee
Abstract
Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts. However, current MLLMs remain fundamentally limited by object hallucination-generating plausible yet factually inconsistent descriptions about objects. Existing benchmarks, designed primarily for single-image settings or providing only high-level multi-image assessments, cannot systematically diagnose how visual complexity and reasoning demands trigger hallucination. To address this gap, we introduce MIOH, a fine-grained multi-image object hallucination benchmark that systematically evaluates object hallucination across four foundational tasks (existence, counting, attribute, position) through three multi-image reasoning patterns (comprehensive, comparative, selective) under three controlled adversarial pressures (visual context scale, perceptual difficulty, contextual bias). Through evaluation of 29 models, we reveal that even state-of-the-art systems like GPT-5 and Gemini-2.5-Pro exhibit distinct failure patterns across different reasoning patterns and tasks. Our evaluation reveals that hallucination stems not merely from perceptual failures but from integration-stage limitations when maintaining object representations across multiple images. MIOH provides a controlled framework for analyzing multi-image object hallucination and serves as a critical evaluation tool for developing more reliable multimodal AI systems.