Generative uncertainty improves video semantic similarity learning
Generative Uncertainty as a Self-supervised Signal for Semantic Similarity Learning
Computer Vision and Pattern Recognition
Summary
Figuring out how similar two videos are is really hard because videos have lots of moving parts and details. The authors found that when AI models try to create videos from text, they make very similar videos for common ideas but get confused and uncertain for unusual ones. By using this uncertainty as a signal, the authors taught another AI model to focus on the stable, important parts of video features and ignore noisy or uncertain parts. This method lets the AI understand video similarity better without needing people to label lots of videos.
What this means in practice
- •For video search engineers: Enhance video retrieval systems by using learned stable features that better measure video semantic similarity without manual labeling.
- •For security analytics teams: Improve detection of unusual or out-of-distribution video content by focusing on uncertainty signals from video generative models.
Authors
Enrico Pallotta, Sina Raoufi, Lars Doorenbos, Gianni Franchi, Juergen Gall
Abstract
Evaluating semantic similarity between videos is a fundamental challenge in computer vision, essential for tasks ranging from out-of-distribution (OOD) detection to video retrieval. However, defining and labeling video similarity is notoriously difficult and expensive due to the complex spatio-temporal nature. In this paper, we propose a novel self-supervised approach that leverages generative uncertainty from text-to-video (T2V) diffusion models to learn semantic similarity without human annotations. Our method is based on the observation that T2V models produce consistent outputs for familiar concepts but exhibit high variance and uncertainty when prompted with specialized concepts. We utilize this behavior to identify stable semantic features within existing pretrained representations, such as VideoMAE and V-JEPA. Specifically, we learn a mask over these embeddings using purely generated data, encouraging the model to retain features that remain consistent across generations of general concepts while discarding those associated with generative noise or uncertainty. Experimental results across three key tasks demonstrate that our learned feature subspaces consistently outperform original pretrained features and baseline feature selection methods.