Omnimodal agent improves referring video segmentation with precise reasoning

OPERA: A Unified Omnimodal Progressive Spatio-Temporal Reasoning Agent for Referring Video Segmentation

Computer Vision and Pattern Recognition

Summary

Referring video segmentation means identifying specific things or moments in a video based on clues like words, sounds, or pictures. The authors created OPERA, a system that carefully looks through video frames over time and uses different types of clues to find exactly what to highlight in the video. OPERA filters information step-by-step to focus on important frames and spots, producing detailed masks that outline the target. This new system works better than previous ones on several video segmentation tests.

What this means in practice

  • For video editing teams: Automatically highlight and segment targets in videos based on text, audio, or images to streamline editing tasks.
  • For security monitoring operators: Identify and track objects or events in surveillance videos more precisely using multiple forms of input like audio and images.

Authors

Jingchen Ni, Yuji Wang, Shannan Yan, Haoru Li, Sitong Chen, Chun Yuan

Abstract

Referring video segmentation with heterogeneous multimodal queries---spanning text, audio, and reference images---demands both robust cross-modal understanding and precise spatio-temporal reasoning. We propose OPERA (Omnimodal Progressive spatio-tEmporal Reasoning Agent), a unified reasoning agent built on a single MLLM that performs dual-axis progressive reasoning via three specialized stages. Along the temporal axis, a Temporal Reasoning Agent narrows the frame search space through coarse-to-fine filtering to identify the most informative key frame. Along the spatial axis, a Distillation Agent establishes what to locate via cross-modal semantic distillation, and a Grounding Agent enhanced with GRPO determines where the target appears, with dense mask propagation completing the pixel-level output. OPERA sets a new state of the art on OmniAVS and Ref-AVS and transfers zero-shot to standard referring video segmentation benchmarks.