Modality aware retrieval improves multimodal question answering performance

Signal or Noise? Modality Contribution and Cooperation in Multimodal GraphRAG

Information Retrieval

Summary

When computers answer questions by looking at documents with text, tables, and pictures, it is often assumed that using all these parts helps. The authors studied whether getting information from every type of content actually helps or just adds confusing or repeated facts. They found that text and tables are usually the most helpful, and adding more modalities often brings redundancy, especially when text is involved. Sometimes, non-text types like images and tables work well together, but it depends on the question and task. This means it is better to pick which content types to use for each question, rather than using everything.

What this means in practice

  • For document understanding teams: Improve question answering systems by selectively retrieving text and table information instead of all modalities to reduce distracting and redundant evidence.
  • For multimodal ai system developers: Design knowledge graph retrieval that filters modalities based on question type to enhance inference accuracy in complex multimodal tasks.

Authors

Antonios Georgakopoulos, Paul Groth, Lise Stork

Abstract

Multimodal knowledge graphs (KGs) integrate information from text, figures, tables, and other modalities into a unified structured representation, with the promise that richer evidence enables better inference. In GraphRAG systems built over such graphs, it is commonly assumed that retrieving evidence from more modalities at inference time improves downstream performance. Yet, redundant or overlapping multimodal evidence may distract language models in question answering (QA), and whether each modality contributes equally across questions, models, and tasks remains poorly understood. In this work, we study how modality-aware retrieval affects downstream inference in a multimodal GraphRAG pipeline, using document visual question answering (DocVQA) as a testbed. We extend an existing KG-based QA framework to be modality-aware, leveraging the graph structure to track which modality supports which facts and to selectively filter evidence at the edge level. This enables us to investigate whether providing all available multimodal evidence at inference time benefits QA, and to evaluate the contribution and cooperation of modalities across question, task, and model characteristics. Through a controlled analysis within a state-of-the-art multimodal GraphRAG pipeline, five multimodal LLMs and two DocVQA benchmarks, we find that tables and text provide the strongest contributions, and that combining modalities frequently produces redundancy rather than synergy, particularly for pairs involving textual information. Positive cooperation appears mainly between non-text modalities and depends on question intent and task type. Our findings argue for selective, modality-aware retrieval in the design of more effective GraphRAG systems, where modalities are filtered according to the downstream task rather than retrieved uniformly.