Papers for

multimodal ai system developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Modality aware retrieval improves multimodal question answering performance

Signal or Noise? Modality Contribution and Cooperation in Multimodal GraphRAG

Abstract: Multimodal knowledge graphs (KGs) integrate information from text, figures, tables, and other modalities into a unified structured representation, with the promise that richer evidence enables better inference. In GraphRAG systems built over such graphs, it is commonly assumed that retrieving evidence from more modalities at inference time improves downstream performance. Yet, redundant or overlapping multimodal evidence may distract language models in question answering (QA), and whether each modality contributes equally across questions, models, and tasks remains poorly understood. In this work, we study how modality-aware retrieval affects downstream inference in a multimodal GraphRAG pipeline, using document visual question answering (DocVQA) as a testbed. We extend an existing KG-based QA framework to be modality-aware, leveraging the graph structure to track which modality supports which facts and to selectively filter evidence at the edge level. This enables us to investigate whether providing all available multimodal evidence at inference time benefits QA, and to evaluate the contribution and cooperation of modalities across question, task, and model characteristics. Through a controlled analysis within a state-of-the-art multimodal GraphRAG pipeline, five multimodal LLMs and two DocVQA benchmarks, we find that tables and text provide the strongest contributions, and that combining modalities frequently produces redundancy rather than synergy, particularly for pairs involving textual information. Positive cooperation appears mainly between non-text modalities and depends on question intent and task type. Our findings argue for selective, modality-aware retrieval in the design of more effective GraphRAG systems, where modalities are filtered according to the downstream task rather than retrieved uniformly.

Mon 28 SeptInformation Retrieval
The gist
When computers answer questions by looking at documents with text, tables, and pictures, it is often assumed that using all these parts helps. The authors studied whether getting information from every type of content actually helps or just adds confusing or repeated facts. They found that text and tables are usually the most helpful, and adding more modalities often brings redundancy, especially when text is involved. Sometimes, non-text types like images and tables work well together, but it depends on the question and task. This means it is better to pick which content types to use for each question, rather than using everything.
Open → 2609.35304v1

Joint audio video generation improves fidelity and sync with new RL method

AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

Abstract: Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited per-modality fidelity, insufficient text-modality alignment and weak cross-modal synchronization. While reinforcement-learning post-training offers a promising remedy, directly adapting it to joint audio-video generation is challenging. Heterogeneous multimodal rewards entangle learning signals and complicate credit assignment. Joint optimization of two modality towers is computationally expensive given their divergent dynamics. Moreover, synchronization evaluation difficulty depends on paired samples, preventing fair reward comparisons. We propose AV-GRPO, a modality-anchored online diffusion RL framework, and 5DAV, a decoupled, difficulty-controllable training dataset. AV-GRPO includes three key modules: (1) modality-anchored rollouts to disentangle learning signals and stabilize difficulty; (2) trajectory-locked frozen-tower optimization to reduce cost and reassign credit; (3) adaptive objectives and perturbation strengths tailored to modality-specific dynamics. This converts coupled multimodal preference learning into unimodal subproblems for precise reward attribution and better synchronization. Our 5DAV dataset decouples samples across five dimensions for systematic training. Experiments on JavisBench and VABench demonstrate AV-GRPO outperforms LTX-2.3 in generation quality, semantic alignment and cross-modal synchronization under LoRA and full fine-tuning. Ablations confirm our designs. Code and data: https://github.com/zhiyuxu03/AV-GRPO

Thu 24 SeptComputer Vision and Pattern RecognitionSound
The gist
Generating videos with matching sounds is hard because models struggle to make each part realistic and well-timed together. The researchers created AV-GRPO, a new learning approach that separates audio and video training to give clearer feedback and better timing. They also made 5DAV, a special dataset to help train models more effectively. Tests showed their method produces better quality videos and sounds that match well with the text descriptions and each other.
Open → 2609.29816v1