Papers for

document understanding teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Modality aware retrieval improves multimodal question answering performance

Signal or Noise? Modality Contribution and Cooperation in Multimodal GraphRAG

Abstract: Multimodal knowledge graphs (KGs) integrate information from text, figures, tables, and other modalities into a unified structured representation, with the promise that richer evidence enables better inference. In GraphRAG systems built over such graphs, it is commonly assumed that retrieving evidence from more modalities at inference time improves downstream performance. Yet, redundant or overlapping multimodal evidence may distract language models in question answering (QA), and whether each modality contributes equally across questions, models, and tasks remains poorly understood. In this work, we study how modality-aware retrieval affects downstream inference in a multimodal GraphRAG pipeline, using document visual question answering (DocVQA) as a testbed. We extend an existing KG-based QA framework to be modality-aware, leveraging the graph structure to track which modality supports which facts and to selectively filter evidence at the edge level. This enables us to investigate whether providing all available multimodal evidence at inference time benefits QA, and to evaluate the contribution and cooperation of modalities across question, task, and model characteristics. Through a controlled analysis within a state-of-the-art multimodal GraphRAG pipeline, five multimodal LLMs and two DocVQA benchmarks, we find that tables and text provide the strongest contributions, and that combining modalities frequently produces redundancy rather than synergy, particularly for pairs involving textual information. Positive cooperation appears mainly between non-text modalities and depends on question intent and task type. Our findings argue for selective, modality-aware retrieval in the design of more effective GraphRAG systems, where modalities are filtered according to the downstream task rather than retrieved uniformly.

Mon 28 SeptInformation Retrieval
The gist
When computers answer questions by looking at documents with text, tables, and pictures, it is often assumed that using all these parts helps. The authors studied whether getting information from every type of content actually helps or just adds confusing or repeated facts. They found that text and tables are usually the most helpful, and adding more modalities often brings redundancy, especially when text is involved. Sometimes, non-text types like images and tables work well together, but it depends on the question and task. This means it is better to pick which content types to use for each question, rather than using everything.
Open → 2609.35304v1