Video text models improve by adding diverse detailed captions
Advancing Video-Text Pretraining with Multi-View Captions
Computer Vision and Pattern Recognition
Summary
Video-text models learn to connect videos with their descriptions, but many existing datasets only have one short caption per video, which misses important details. The authors created a method to generate multiple kinds of captions for each video, including summaries and detailed descriptions, to better cover what happens in the video. They also designed a way for the model to understand these different caption types separately. Models trained on these richer captions performed better at retrieving videos from text queries, even when tested without extra training. This shows that improving the quality and diversity of captions helps models understand videos better.
What this means in practice
- •For video search developers: Enhance video search engines by training on diverse detailed captions to improve matching of complex video content to text queries.
- •For media content analysts: Improve automated video indexing with multi-view captions to capture detailed and summary information for better content understanding.
Authors
Fida M. Thoker, Renaud Vandeghen, Karen Sanchez, Marc Van Droogenbroeck, Bernard Ghanem
Abstract
Video-text pretraining has achieved remarkable progress through the scaling of models and datasets, yet the quality of language supervision remains underexplored. Existing web-scale datasets often provide only a single sparse caption per video that fails to capture rich spatiotemporal semantics, while directly using captioning models can generate noisy descriptions. We propose a large-scale multimodal large language model-based supervision generation framework that improves supervision diversity, fidelity, and semantic coverage. Starting from 10 million videos, our approach generates multi-view captions (MVC) through complementary summary and detailed captions, reasoning-based refinement, and semantic positive caption generation. To effectively exploit supervision at different granularities, we further introduce a granularity-aware text representation with separate CLS tokens for summary and detailed views. We pretrain video-text models using the resulting supervision corpus and evaluate them across standard, fine-grained and detailed text-to-video retrieval benchmarks. Our approach consistently improves both zero-shot and fine-tuned performance while using smaller pretraining corpora than existing methods, demonstrating the importance of rich and complementary textual supervision for video-text pretraining. Project page: https://rvandeghen.github.io/mvc/