OmniSeek improves multi-turn audio visual reasoning with built in tool use
OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning
Computer Vision and Pattern Recognition
Summary
Understanding important details in long videos or audio clips can be hard because there is too much information. The authors created OmniSeek, a smart system that actively chooses when and where to look or listen to find key clues before making decisions. It learns by practicing on lots of example reasoning steps involving both sounds and images and gets better through a special kind of trial-and-error training. OmniSeek can combine audio and visual clues over several turns to make better judgments than if it looked at everything at once.
What this means in practice
- •For multimedia content analysts: Automatically gather key visual and audio segments from long multimedia files to improve multi-step understanding and decision-making.
- •For digital assistant developers: Enable assistants to actively seek and integrate audio-visual information over multiple interactions, enhancing multimedia question answering.
Authors
Haibo Wang, Jiteng Mu, Jialu Li, Jingru Yi, Yuanjun Xiong, Jianming Zhang, Lifu Huang, Mingze Xu
Abstract
We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with native tool use. Rather than passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence acquisition part of the reasoning process: it dynamically decides whether to look or listen, and over which temporal window, to retrieve sparse but critical evidence across different modalities within long contexts. Through an iterative multi-turn protocol, the retrieved raw audio or visual segments are appended back into the context to support subsequent reasoning. To cold-start this capability, we build a data engine that synthesizes OmniTraj-170K, a corpus of multi-hop Chain-of-Thought trajectories with interleaved audio and visual evidence. We first supervise the model on these trajectories to instill multi-turn tool-use behavior, and then further optimize the policy via a two-stage reinforcement learning with verifiable rewards. Moreover, we introduce an Audio-Visual Necessity objective that explicitly rewards successful trajectories whose reasoning depends on both modalities, discouraging single-modality shortcuts. Extensive experiments across a wide range of benchmarks demonstrate that OmniSeek learns adaptive cross-modal evidence seeking and consistently improves audio-visual reasoning performance.