Papers for

multimodal system integrators

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Multimodal agent improves retrieval by folding context and images

MM-ContextFold: Context Folding for Multimodal Agentic Retrieval

Abstract: Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical frameworks such as ReAct maintain raw multimodal inputs and the accumulating interaction history in a single, ever-growing context, leading to the context explosion problem. While existing methods alleviate this issue by compressing redundant text, effective strategies for managing token-intensive visual content remain largely underexplored. To address this gap, we first conduct a systematic empirical study of approximately 10,000 trajectories. The results show that as visual cues are progressively extracted through external tools and textualized into the context, raw images become increasingly redundant. Continued image retention is associated with higher output entropy and can even degrade task accuracy. Motivated by these findings, we propose MM-ContextFold, a training-free framework that loads raw images only when needed. It maintains a persistent, text-only main context for high-level planning and spawns ephemeral branch contexts for image-dependent subtasks. Within each branch, the agent loads the relevant images, completes the subtask, and folds the result back into the main context as a concise textual summary; the images and branch trace are then discarded. Experiments on seven MAR benchmarks across five backbone models show that MM-ContextFold improves average accuracy by 6.3 percentage points over ReAct while reducing the working context length by 27.5\%.

Sat 19 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceInformation Retrieval
The gist
Searching for information using AI agents gets tricky when they have to handle both pictures and text, as keeping all that data in one place becomes too large and confusing. The authors studied many cases and found that keeping raw images after turning them into text doesn’t help and can even cause mistakes. They created MM-ContextFold, a method that only loads images when necessary and summarizes what’s learned into text, keeping the main workspace smaller and cleaner. This approach made the agents more accurate and efficient in seven different tasks.
Open → 2609.23121v1

Visual grounding improves with confidence awareness to reduce hallucinations

SAVOR: Self-Aware Visual Grounding via Confidence-Calibrated Reinforcement Learning for Multimodal Hallucination Mitigation

Abstract: Multimodal large language models (MLLMs) have made strong progress on visual question answering and image captioning, yet they still produce fluent claims about objects, attributes, or relations that are not grounded in the image. Many remedies either modify decoding at test time, which adds latency, or fine tune with preferences such as DPO variants, which teach which answer is preferred but not when the model's own answer is unreliable. We argue that calibrated self assessment is the missing signal. We introduce Savor, a training framework that (i) augments the output schema with token and answer confidence, (ii) optimises the policy with a Group Relative Policy Optimisation (GRPO) objective that penalises calibration error and poor abstention decisions, and (iii) uses the learned confidence at inference time to revisit visual evidence only when the model is uncertain. Experiments on POPE, HallusionBench, AMBER and MMHal-Bench across two recent backbones (InternVL3-8B and Qwen3-VL-8B) show that Savor reduces hallucination while preserving general capability on MME and MMBench, with lower Expected Calibration Error than DPO and decoding baselines.

Tue 15 SeptComputer Vision and Pattern Recognition
The gist
Multimodal large language models sometimes make confident but incorrect claims about images, describing things that aren’t really there. The authors found that teaching these models to be aware of their own confidence helps them avoid making false statements. They created a training method called Savor that lets the model judge when it is uncertain and look again at the image before answering. This approach reduces mistakes while keeping the model’s overall ability to understand images and text.
Open → 2609.16601v1