Papers for

video content indexing teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Audio visual models mistakenly match speech to faces by position

Uncovering Ordinal-Matching Bias in Audio-Visual LLMs

Abstract: This work aims to improve how audio-visual large language models (AVLLMs) associate speech with the correct visible speaker in multi-speaker scenes. We find that current AVLLMs frequently fail at this task, and analyze the nature of these failures. To this end, we construct a synthetic diagnostic dataset in which multiple visible speakers each utter a single word. Analysis on this corpus reveals a consistent error pattern across three recent open-source AVLLMs: models attribute utterances by simply matching the order of spoken sentences with the left-to-right, top-to-bottom arrangement of visible faces, rather than relying on audio-visual cues such as lip synchronization. We term this behavior \emph{ordinal-matching bias}. We further show that this bias can be substantially mitigated through a simple remedy, Ordinal-Decoupled Fine-Tuning (OD-FT), in which models are fine-tuned on synthetic videos where spatial positions of speakers and speaking order are independently randomized. Despite using only 400 synthetic training videos, OD-FT not only suppresses ordinal-matching bias but also improves audio-visual understanding on real-world videos, yielding average gains of 8.27\% for Qwen2.5-Omni and 2.57\% for video-SALMONN2+ across three audio-visual benchmarks.

Mon 28 SeptComputer Vision and Pattern RecognitionSound
The gist
Audio-visual large language models often try to match spoken words to faces in videos, but the authors found they usually just pair speech to the face's position on screen instead of using real audio and visual cues. They call this mistake ordinal-matching bias. To fix this, the authors created a new training method that randomizes speaker positions and speaking order during fine-tuning, which helps the models learn to associate speech and faces correctly. This improvement makes the models better at understanding who’s speaking in real videos.
Open → 2609.34223v1

Dual-side enhancements improve video segment retrieval without training

DSE-VTG: Dual-Side Enhancement for Training-Free Video Temporal Grounding

Abstract: Text-guided Video Temporal Grounding (VTG) aims to localize the relevant segments in an untrimmed video based on text queries, yet collecting dense temporal annotations and training task-specific models remain costly and brittle under distribution shift. Recent training-free VTG approaches mitigate this issue by directly matching pretrained vision-language representations, but they still face two fundamental information bottlenecks: frame-wise visual encoding overlooks temporal dynamics, while fixed query embeddings cannot resolve query ambiguity. To address these issues, we propose DSE-VTG, a \underline{D}ual-\underline{S}ide \underline{E}nhancement framework that addresses both without any task-specific training. On the visual side, Multi-scale Similarity Fusion (MSF) combines frame- and clip-level similarities into a unified, temporally aware similarity profile. On the textual side, Query-level Test-Time Adaptation (Q-TTA) optimizes a lightweight additive offset to adapt the query embedding to the video at test time, without finetuning the backbone or calling external large language models. Extensive experiments on three standard and two OOD benchmarks show that DSE-VTG achieves state-of-the-art performance among training-free methods. On Charades-STA, it improves mIoU over the strongest prior training-free method by 5.61 points. Under distribution shift, DSE-VTG reaches 50.86 mIoU on Charades-CG Novel-Word, surpassing the strongest supervised baseline by 2.76 mIoU. Our code will be released upon acceptance.

Tue 8 SeptComputer Vision and Pattern Recognition
The gist
Finding the right part of a long video that matches a text description is hard and usually needs a lot of training. The paper introduces a new method that doesn’t require any extra training but still gets better at this task. They improve how videos are compared by looking at both single frames and groups of frames, and they also tweak the text description to better fit the video’s content. This approach works well even when the videos are different from the training examples used by other methods.
Open → 2609.08850v1