Papers for

multimodal system engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

VideoLoop improves long video understanding by managing memory better

VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents

Abstract: Long-form video understanding requires multimodal agents to iteratively gather evidence over many reasoning steps. However, most existing agentic methods suffer from semantic thrashing: as append-only working memory grows, attention to key evidence collapses, and the agent loses access to what it has already found. First, we provide a structural argument showing that append-only memory can incorporate newly observed target evidence, but cannot remove accumulated noise or prevent ordered context growth without a rewrite operator. Second, motivated by this analysis, we propose VideoLoop, a multimodal agent with two coupled loops. The outer loop reasons over the video and the inner loop, after each step, retrieves artifacts from an unbounded filesystem of past observations and intermediate analysis, and rewrites a bounded working memory. Extensive experiments demonstrate the effectiveness of VideoLoop, which improves four popular LVLM backbones in a plug-and-play manner, with an average gain of 4.2% points over baseline on VideoMME (long). Further analysis of working memory suggests that VideoLoop mitigates semantic thrashing: on the hardest quarter of VideoMME (long) questions, a blind judge that reads only the agent's context answers 81.1% correctly, versus 60.9% for the append-only agent. With Gemini 3.1 Pro, VideoLoop reaches 88.3% on VideoMME (long), 88.8% on VideoMMMU, and 80.9% on LongVideoBench (long).

Tue 29 SeptComputer Vision and Pattern Recognition
The gist
Understanding long videos needs AI systems to think step-by-step, remembering important details while ignoring distractions. The authors found that simply adding more memory can cause the system to get confused by unrelated information. They created VideoLoop, which uses two memory loops to keep track of key points and rewrite its memory to stay focused. This approach helps AI agents answer questions about long videos more accurately than before.
Open → 2609.38119v1

Multimodal models struggle to understand user demands in interactions

Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

Abstract: Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental question underexplored: can a model correctly infer the user's underlying demand from complex multimodal interaction? Real-world user demands are often underspecified in speech and must be inferred from multimodal cues and dialogue history. This inference is further complicated by ambiguous or disfluent expression and noisy acoustic environments. Conversely, request-like speech may not constitute a demand to the assistant, leading to false triggers. We establish Omni Demand Understanding (ODU) as a distinct multimodal contextual inference problem: given an interaction stream, a model must detect whether a user demand is present and infer intent from multimodal and conversational context. ODU evaluates this capability along five dimensions, covering both single-turn and multi-turn interactions. We construct ODU-Bench using a challenge-driven taxonomy, taxonomy-guided agentic video generation, and human-recorded interactions, followed by media-grounded annotation and human verification. We evaluate 14 native MLLMs. Even the strongest, Gemini 3.1 Pro, recovers only 44.7% of key information that must be inferred from visual, acoustic, or conversational context. Moreover, 11 of the 14 models exhibit false-trigger rates above 50% on non-demand scenarios. These results reveal a systematic capability gap in current MLLMs' ability to infer contextual user demands. We hope ODU can establish the evaluation of a previously underexplored yet essential capability in multimodal interaction: correctly understanding user demands before generating an appropriate response.

Fri 18 SeptComputation and LanguageComputer Vision and Pattern RecognitionMultimedia
The gist
People often talk and use gestures or visual cues when asking for help from AI assistants, but these requests can be unclear or noisy. The paper’s authors introduce a new test called Omni Demand Understanding to see if AI systems can figure out what users really want from mixed speech, visuals, and conversation history. They found that even some of the best AI models miss a lot of important details and often think someone is asking for help when they are not. This shows that current AI assistants have trouble truly understanding what people want during natural, everyday conversations.
Open → 2609.21392v1

Unified multimodal models improve image generation and understanding together

Dreaming in Flow: Generative Grounding Feedback for Self-Evolving Unified Multimodal Models

Abstract: Unified multimodal models integrate visual understanding and generation within a single network, yet the two capabilities are commonly optimized as separate tasks. We introduce Generative Grounding Feedback(GGF), a self-evolving post-training framework that uses only text prompts and the model's own visual experience. Given a prompt, the model first generates a visual ``dream.'' Flow-level feedback compares text-, image-, and repair-conditioned predictions at the same noisy latent state, transferring image-grounded generation directions to the prompt condition. Dream replay grounding replays this dream through captioning and re-imagination, training claim-level evidence to remain consistent across the replay while separating unrelated visual experiences. Jointly optimized, these two directions let generation provide visual grounding for understanding and understanding refine subsequent generation without paired image--text supervision. Experiments across unified models with different understanding--generation integration designs show consistent improvements in text-to-image generation together with modest gains in visual understanding.

Tue 8 SeptComputer Vision and Pattern Recognition
The gist
Unified multimodal models often handle image understanding and image creation as separate tasks. The authors introduce a way for these models to use their own generated images to get better at both tasks at the same time without needing special paired training data. Their approach involves the model 'dreaming' a picture from text, checking that picture’s fit with the text, and then replaying this process to improve consistency. This self-feedback method helps the model get better at creating images from text and also modestly improves its ability to understand images.
Open → 2609.08282v1