Papers for

video editing teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Omnimodal agent improves referring video segmentation with precise reasoning

OPERA: A Unified Omnimodal Progressive Spatio-Temporal Reasoning Agent for Referring Video Segmentation

Abstract: Referring video segmentation with heterogeneous multimodal queries---spanning text, audio, and reference images---demands both robust cross-modal understanding and precise spatio-temporal reasoning. We propose OPERA (Omnimodal Progressive spatio-tEmporal Reasoning Agent), a unified reasoning agent built on a single MLLM that performs dual-axis progressive reasoning via three specialized stages. Along the temporal axis, a Temporal Reasoning Agent narrows the frame search space through coarse-to-fine filtering to identify the most informative key frame. Along the spatial axis, a Distillation Agent establishes what to locate via cross-modal semantic distillation, and a Grounding Agent enhanced with GRPO determines where the target appears, with dense mask propagation completing the pixel-level output. OPERA sets a new state of the art on OmniAVS and Ref-AVS and transfers zero-shot to standard referring video segmentation benchmarks.

Sun 27 SeptComputer Vision and Pattern Recognition
The gist
Referring video segmentation means identifying specific things or moments in a video based on clues like words, sounds, or pictures. The authors created OPERA, a system that carefully looks through video frames over time and uses different types of clues to find exactly what to highlight in the video. OPERA filters information step-by-step to focus on important frames and spots, producing detailed masks that outline the target. This new system works better than previous ones on several video segmentation tests.
Open → 2609.33338v1

Binaural audio improves spotting sound sources in videos

Binaural Audio-Visual Instance Segmentation

Abstract: Audio-visual segmentation (AVS) aims to segment sounding objects at the pixel level by integrating auditory and visual cues. However, existing methods are predominantly developed under the monaural setting and primarily rely on cross-modal semantic correspondence, which limits their ability to distinguish visually similar instances of the same semantic class. In contrast, humans naturally exploit binaural hearing, where interaural differences and direction-dependent acoustic filtering introduced by the head and pinnae provide physically grounded spatial cues for accurate sound source localization. Motivated by this observation, we introduce binaural audio-visual instance segmentation (BiAVIS), a new task that leverages synchronized binaural audio and video frames to segment sounding instances. To advance research on this task, we establish two benchmarks by manually annotating an existing binaural audio-visual dataset and collecting a new real-world dataset, BiAVIS-Bench, in more challenging and diverse scenarios. We further propose a BiAVIS model, which leverages an audio-only sound source localization network to learn spatial and semantic priors for sounding instances from binaural audio. A query-level audio-visual fusion strategy is subsequently introduced to inject these informative priors into the instance segmentation decoder. Extensive experiments conducted on the two proposed benchmarks demonstrate the superior performance of the BiAVIS model over previous monaural AVS methods, especially in resolving instance-level intra-class ambiguity. On the more challenging BiAVIS-Bench, the proposed BiAVIS model outperforms the best-performing monaural baselines by 17.22\% in mAP and 7.71\% in FSLA, respectively.

Sat 26 SeptComputer Vision and Pattern RecognitionSound
The gist
It's hard for computers to tell apart objects that look very similar but make different sounds just by watching videos with one microphone. The authors help computers by using two microphones, like how human ears work, to figure out where sounds are coming from more precisely. They created new tests and a method that uses these two microphones to better find and separate the exact sound-making objects in videos. Their method works better than older ones, especially when objects look alike but sound different.
Open → 2609.32180v1

EviDETR improves video search and highlight spotting from text queries

EviDETR: Preserving Query-Relevant Temporal Evidence for Moment Retrieval and Highlight Detection

Abstract: Joint video moment retrieval and highlight detection requires identifying query-relevant temporal segments while estimating clip-level saliency, yet DETR-style pipelines do not explicitly preserve query-relevant evidence throughout encoding, decoding, and cross-task prediction. We propose EviDETR, an evidence-preserving framework with three components. Semantic-aware Feature Reweighting (SFR) enhances query-relevant clip representations through saliency estimation and cross-modal interaction. A Temporal Top-2 Mixture-of-Experts (TTop2MoE) decoder performs query-adaptive refinement via sparse expert routing. MR-to-HD (MR2HD) fusion transfers span-level retrieval evidence to clip-level highlight prediction through confidence-weighted multi-scale aggregation. Using CLIP+SlowFast features, EviDETR achieves 69.29 R1@0.5, 54.77 R1@0.7, and 48.41 Avg. mAP for moment retrieval on QVHighlights, together with 41.83 HD-mAP and 68.33 HIT@1. Strong results on TACoS and Charades-STA further demonstrate cross-dataset transferability.

Fri 25 SeptComputer Vision and Pattern Recognition
The gist
Finding specific moments in videos and spotting highlights based on a text query is tricky because existing methods don’t always keep track of what parts of the video relate to the query. The authors propose EviDETR, a new system that better remembers relevant video segments during analysis and prediction. It uses special techniques to focus on important clips and combines information effectively to improve both search and highlight detection. EviDETR shows strong performance on several video datasets, making it easier to find and highlight moments in long videos.
Open → 2609.30724v1

Video deraining improved by fluid-inspired motion guidance

FluidRain: Incompressible Rain Flow as an Attention Bias for Loop-in-Loop Video Deraining

Abstract: Existing video deraining methods typically exploit neighboring frames through either explicit alignment or implicit spatiotemporal aggregation. Explicit alignment relies on accurate motion estimation, which can become unreliable under dense rain, while implicit aggregation avoids alignment but lacks explicit guidance on the directional and temporally coherent structure of rain. This leaves a gap between reliable temporal aggregation and explicit modeling of rain motion. To address these limitations, we propose FluidRain, a lightweight video derainer that uses divergence-free rain flow to guide Loop-in-Loop attention across scales and neighboring frames. Motivated by fluid mechanics, we model rain motion as a divergence-free image-space flow and use it to organize multi-scale and temporal aggregation. Specifically, FluidRain first estimates a rain-flow field for each frame and projects it onto the divergence-free subspace. The resulting flow steers window attention along rain streaks, enabling neighboring frames to be aggregated without explicit alignment. Since rain-flow structure is preserved across scales and nearby frames, Loop-in-Loop reuses the same attention operator across both dimensions, resulting in a three-frame model with only 0.80M parameters. Experiments on four benchmarks show that FluidRain remains competitive with substantially larger restoration models. We further examine how temporal evidence scales with different input views. To evaluate whether the model remains reliable when rain motion changes across frames, we introduce RainSyn-Gust, which injects controlled changes in rain-streak direction into existing benchmarks. We also develop a physics-based no-reference metric that evaluates real-rain removal without requiring clean targets.

Thu 24 SeptComputer Vision and Pattern Recognition
The gist
Rain in videos can make it hard to see clearly, and removing rain usually involves matching or blending nearby video frames. Existing methods either rely on matching moving parts, which is hard when rain is thick, or blend frames without clear direction, missing the rain’s motion pattern. The authors created FluidRain, a method that models rain movement inspired by fluid mechanics, guiding frame blending along rain streaks without complex matching. This approach uses only a small, efficient model that works well across different videos and rain conditions.
Open → 2609.29006v1

Large AI models struggle to understand cinematic storytelling techniques

CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models

Abstract: Cinematography, the craft of visual storytelling through framing, lighting, and camera operation, fundamentally shapes how audiences perceive and emotionally engage with video content. While Large Vision Language Models (LVLMs) have made remarkable progress in video question answering, existing benchmarks primarily focus on identifying low-level techniques rather than understanding their storytelling impact. To address this, we introduce CinematicVQA, the first-of-its-kind benchmark for cinematic video understanding that goes beyond technique recognition to evaluate film-grammar reasoning, utilizing our introduced Cinematic Scene Graph (CSG), a structured representation that links filming techniques to their perceptual effects and narrative functions. Through comprehensive evaluation of state-of-the-art LVLMs, we reveal a striking semantic gap: models consistently perform higher on describing visual presentations than on identifying the underlying techniques. Surprisingly, Chain-of-Thought prompting fails to provide consistent gains and degrades performance for most models, suggesting that current LVLMs lack sufficient cinematic domain knowledge to benefit from step-by-step reasoning. Fine-tuning on \textsc{CinematicVQA-train} yields consistent improvements, particularly for narrative function and multi-hop reasoning. Overall, \textsc{CinematicVQA} serves both as a rigorous benchmark for cinematic evaluation in LVLMs and as a practical dataset for training more film-aware video models.

Wed 23 SeptComputer Vision and Pattern Recognition
The gist
Understanding how movies tell stories visually is hard for AI models that see and read videos. The authors created a new test called CinematicVQA, which checks if AI can reason about the film-making methods behind what we see, not just describe the images. They found that AI models are better at describing scenes visually than explaining the storytelling tricks used, and common techniques to improve reasoning did not help. Training AI on this new test helped them get better at understanding narrative and multi-step story reasoning.
Open → 2609.28813v1

Video prompt inversion benchmark reveals current model limitations

VI-Bench: Benchmarking Prompt Inversion from AIGC Videos

Abstract: Recent advances in video generation have made prompt-based control increasingly central to AIGC video generation. Prompts specify what a video should depict and how it should be represented, controlling factors such as visual style or camera behavior. Understanding this recoverability is important both for creative reuse and editing, and for assessing prompt leakage risks. However, existing video understanding benchmarks do not measure this capability: a caption may describe what is visible, but a replayable prompt must recover the generation-relevant controls needed to reproduce the video. To address this gap, we introduce VI-Bench, a benchmark built from 16.1 million real-user prompts and 900 human-verified AIGC videos. VI-Bench spans three progressively harder settings, namely single-shot semantic grounding, control over style and camera behavior, and multi-shot compositional inversion, and evaluates five generation-critical dimensions: subject, action, scene, style, and camera. We evaluate 18 representative VLMs, including 2 proprietary and 16 open-source models on VI-Bench, using an Inversion Score that measures prompt-level alignment with the original prompt and video-level fidelity of the regenerated video. The results reveal substantial limitations: even the strongest model achieves only 0.632 on Inversion Score, performance degrades sharply as samples require richer control and multi-shot reasoning, and models often produce plausible prompts whose regenerated videos deviate from the reference. These findings show that video prompt inversion is a distinct and under-evaluated capability requiring models to transform visual understanding into replay-stable generative control.

Tue 8 SeptComputer Vision and Pattern Recognition
The gist
Understanding the exact text prompts used to create AI-generated videos is important for editing and reusing those videos. Existing tests only check if a caption describes a video, but don’t measure if the key controls or instructions can be recovered to recreate it. The authors created a new large benchmark named VI-Bench that tests how well different AI models can reverse-engineer these prompts from videos. Their results show that even the best models struggle, especially with complex videos needing detailed control.
Open → 2609.08079v1