Multimodal embeddings read evidence boundaries to improve retrieval accuracy

Learning Multimodal Embeddings with Evidence-Aligned Readout

Computer Vision and Pattern RecognitionInformation Retrieval

Summary

When computers process information from different sources like text and images, they need a way to organize and understand that evidence to find the most relevant answers. The authors created a method called EviAlign that groups related pieces of evidence into meaningful chunks and reads them at specific boundaries to create a better combined understanding. They found that organizing the evidence consistently at these boundaries helps retrieval performance more than just generating evidence. Their approach improves how well the system can find the right information across many tasks, using a single combined representation.

What this means in practice

  • For multimodal ai developers: Create retrieval systems that better combine evidence from text and images for improved task-relevant search results.
  • For enterprise search teams: Improve search accuracy by structuring evidence into semantically aligned chunks for more precise retrieval embeddings.

Authors

Zirong Chen, Fuda Ye, Enjun Du, Junfu Pu, Xinlei Wang, Xinyu Zuo, Lisheng Duan, Haijin Liang, Jin Ma, Jiachuan Wang, Yongqi Zhang

Abstract

Multimodal large language models can expose task-relevant evidence through generation, but producing useful evidence does not by itself determine how it enters a retrieval embedding. We study whether the semantic organization of that evidence can also specify where representations are read. To address this question, we introduce EviAlign, which couples Semantic Evidence Generation with Boundary Readout in a shared multimodal large language model. It organizes evidence into five semantic units, reads the contextualized state at each unit boundary, and aggregates these states into a single normalized embedding. Generation and contrastive retrieval objectives jointly train this shared structure. With the same trailing readout, semantic evidence and free-form CoT yield nearly identical retrieval performance, suggesting that evidence organization alone does not explain the full gain. A controlled $2\times3$ study compares consistent and permuted evidence organization across three readout strategies, using training targets with matched evidence spans. With five readout states and the same mean pooling, the advantage of consistent semantic organization grows from 0.65 points at length-based training positions to 2.39 at evidence boundaries, yielding a 1.74-point co-design interaction. Across 12 MMEB retrieval tasks, EviAlign achieves 76.9 average Recall@1 with 500K training pairs while retaining single-vector indexing and scoring.