Papers for

media content analysts

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Video text models improve by adding diverse detailed captions

Advancing Video-Text Pretraining with Multi-View Captions

Abstract: Video-text pretraining has achieved remarkable progress through the scaling of models and datasets, yet the quality of language supervision remains underexplored. Existing web-scale datasets often provide only a single sparse caption per video that fails to capture rich spatiotemporal semantics, while directly using captioning models can generate noisy descriptions. We propose a large-scale multimodal large language model-based supervision generation framework that improves supervision diversity, fidelity, and semantic coverage. Starting from 10 million videos, our approach generates multi-view captions (MVC) through complementary summary and detailed captions, reasoning-based refinement, and semantic positive caption generation. To effectively exploit supervision at different granularities, we further introduce a granularity-aware text representation with separate CLS tokens for summary and detailed views. We pretrain video-text models using the resulting supervision corpus and evaluate them across standard, fine-grained and detailed text-to-video retrieval benchmarks. Our approach consistently improves both zero-shot and fine-tuned performance while using smaller pretraining corpora than existing methods, demonstrating the importance of rich and complementary textual supervision for video-text pretraining. Project page: https://rvandeghen.github.io/mvc/

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Video-text models learn to connect videos with their descriptions, but many existing datasets only have one short caption per video, which misses important details. The authors created a method to generate multiple kinds of captions for each video, including summaries and detailed descriptions, to better cover what happens in the video. They also designed a way for the model to understand these different caption types separately. Models trained on these richer captions performed better at retrieving videos from text queries, even when tested without extra training. This shows that improving the quality and diversity of captions helps models understand videos better.
Open → 2609.35090v1