Multimodal agent improves retrieval by folding context and images
MM-ContextFold: Context Folding for Multimodal Agentic Retrieval
Computer Vision and Pattern RecognitionArtificial IntelligenceInformation RetrievalMultimedia
Summary
Searching for information using AI agents gets tricky when they have to handle both pictures and text, as keeping all that data in one place becomes too large and confusing. The authors studied many cases and found that keeping raw images after turning them into text doesn’t help and can even cause mistakes. They created MM-ContextFold, a method that only loads images when necessary and summarizes what’s learned into text, keeping the main workspace smaller and cleaner. This approach made the agents more accurate and efficient in seven different tasks.
What this means in practice
- •For ai software engineers: Build information-seeking AI agents that efficiently handle images by loading them only when needed and keeping a concise text context.
- •For multimodal system integrators: Develop retrieval systems that reduce memory use and improve accuracy by folding visual subtasks into summarized textual branches.
Authors
Yang Tian, Fan Liu, Jingyuan Zhang, Zhenyang Li, Yupeng Hu, Liqiang Nie
Abstract
Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical frameworks such as ReAct maintain raw multimodal inputs and the accumulating interaction history in a single, ever-growing context, leading to the context explosion problem. While existing methods alleviate this issue by compressing redundant text, effective strategies for managing token-intensive visual content remain largely underexplored. To address this gap, we first conduct a systematic empirical study of approximately 10,000 trajectories. The results show that as visual cues are progressively extracted through external tools and textualized into the context, raw images become increasingly redundant. Continued image retention is associated with higher output entropy and can even degrade task accuracy. Motivated by these findings, we propose MM-ContextFold, a training-free framework that loads raw images only when needed. It maintains a persistent, text-only main context for high-level planning and spawns ephemeral branch contexts for image-dependent subtasks. Within each branch, the agent loads the relevant images, completes the subtask, and folds the result back into the main context as a concise textual summary; the images and branch trace are then discarded. Experiments on seven MAR benchmarks across five backbone models show that MM-ContextFold improves average accuracy by 6.3 percentage points over ReAct while reducing the working context length by 27.5\%.