Region aware retrieval improves image and text matching in large models
RegRet: Enhancing Region-Level Retrieval in Large Multimodal Models
Computer Vision and Pattern RecognitionArtificial IntelligenceInformation Retrieval
Summary
Finding specific parts of an image and matching them to related words or other image parts is important but hard. The authors designed RegRet, which helps large multimodal models look more closely at image regions without losing the overall picture. They also created a big new dataset to train and test this regional matching. Their experiments show RegRet does a better job than previous methods, especially when it comes to detailed image parts.
What this means in practice
- •For e-commerce platform developers: Improve product search by matching specific parts of product images with detailed textual descriptions or other images.$Commercial implications: Enables enhanced visual search features for online stores, improving user experience and sales through precise product matching.
- •For content management teams: Enhance retrieval of relevant images or image parts in large multimedia databases by better understanding regional details.
Authors
Xun Liang, Honghui Yang, Weihang Pan, Ruisi Zhao, Boyuan Pan, Yao Hu, Wenxiao Wang, Binbin Lin, Deng Cai
Abstract
Region-level retrieval aims to align user-specified image regions with relevant regions or textual descriptions, playing a crucial role in realworld applications such as e-commerce product search and RAG. Although recent Large Multimodal Models (LMMs) have made significant strides in multimodal retrieval, they primarily focus on global-level tasks and struggle to capture effective region-level representations. To bridge this gap, we present RegRet, an LMM-based Region-level Retrieval framework that enhances the regional representations without compromising overall global retrieval performance. At its core, RegRet integrates a Region-Aware Encoder to capture detailed regional features while balancing them with the global background context. To further enhance the fine-grained understanding and discriminability of representations, we design a multi-stage training pipeline that includes detailed localized captioning and regional contrastive learning tasks. In addition, considering the absence of region-level contrastive training data and the limited diversity of evaluation tasks in current benchmarks, we introduce the REGMB benchmark. It comprises 225k contrastive pairs, covering four multimodal retrieval tasks. Extensive experiments validate the effectiveness of our approach. RegRet outperforms strong baselines in the zero-shot setting. Further training with contrastive learning leads to an average improvement of more than 20\% on both REGMB and public benchmarks, while achieving comparable or better results on global-level retrieval tasks.