Video anomaly detection improved with training-free severity probing
Probe-VAD: Ordinal Likelihood Probing for Training-Free Video Anomaly Detection
Computer Vision and Pattern Recognition
Summary
Video anomaly detection tries to find unusual events in long videos without prep work. The authors found that current methods lose important details when turning video into text or limited scores. They created Probe-VAD, which asks a fixed video-language model about levels of anomaly severity using simple yes/no questions. This method keeps subtle differences and ranks anomalies better without extra training, using a clever math step to keep the rankings consistent. Tests show it works well and uses less computing power.
What this means in practice
- •For video surveillance teams: Detect and rank subtle unusual events in security footage without retraining models for each setting.
- •For industrial monitoring teams: Continuously evaluate video streams of machinery to identify and prioritize unusual behaviors with minimal computing effort.
Authors
Jiawei Gu, Qilin Zhao, Tengkuo Guo, Zhiming Zhong, Shuangqing Zhang, Fan Lyu, Fang Zhao, Guo-Sen Xie, Caifeng Shan
Abstract
Video anomaly detection (VAD) aims to localize anomalous events in untrimmed videos. Vision-language models (VLMs) provide rich visual understanding for training-free VAD, but existing approaches impose restrictive interfaces between visual understanding and anomaly scoring. Caption-based pipelines compress visual evidence into text, potentially discarding subtle cues, while direct numerical generation forces the model to express its judgment through a small set of predefined scores. Such interfaces can obscure subtle differences in anomaly severity, causing visually distinct clips to receive similar representations or scores and thereby limiting the resolution of anomaly ranking. We propose \textbf{Probe-VAD}, an ordinal binary-probing framework that directly probes severity preferences from a frozen VLM. Given raw video clips, Probe-VAD queries ten ordered severity thresholds and extracts constrained \textit{YES}/\textit{NO} continuation likelihoods. Their normalized preferences form a cumulative severity profile, from which tail evidence is aggregated into a continuous anomaly score, with isotonic projection enforcing ordinal consistency. Experiments on public VAD benchmarks demonstrate superior performance with low computational cost. Probe-VAD provides a simple interface for translating frozen VLM visual understanding into continuous, rank-sensitive anomaly scores without task-specific training or caption-based compression. Code is available at: https://github.com/yvestine/COVAS-VAD.