Vision language models struggle to link actions correctly over time
ActionLens: Diagnosing Spatial-Temporal Binding Failures in Vision-Language Models
Computer Vision and Pattern RecognitionArtificial IntelligenceComputation and Language
Summary
Many current video understanding AI systems have trouble correctly matching who is doing what and when in videos. The authors created ActionLens, a large set of detailed multiple-choice questions designed to test specific problems in this area, like recognizing different people’s actions and where they are looking. They found that even the best models perform much worse than humans, especially when identifying gaze or matching actions to the right person. The benchmark helps reveal exactly where these models fail so future improvements can be better targeted.
What this means in practice
- •For computer vision developers: Evaluate and diagnose specific action recognition failures in video-based AI models using detailed, human-reviewed questions.
- •For automated video surveillance teams: Improve reliability of detecting who is doing what in complex scenes by using ActionLens to pinpoint model weaknesses.
Authors
Gueter Josmy Faure, Min-Hung Chen, Hao Ping Wang, Timothée Lardy, Hung-Ting Su, Winston H. Hsu
Abstract
Video-capable vision-language models score above 80\% on popular benchmarks yet struggle with spatial-temporal binding: associating the right action with the right person at the right moment. We introduce ActionLens, a diagnostic benchmark of 6,701 multiple-choice video questions spanning five targeted diagnostics: transition detection, actor-specific identification, concurrent action binding, directed interaction reasoning, and gaze detection. Ground-truth answers are derived deterministically from 1.58 million per-second, per-person annotations. Fourteen rounds of human quality engineering raised answer clarity from 53% to above 90% human accuracy. Across 20 VLMs, the full-set leader scores 68.8%; on the human-reviewed subset, it scores 65.9% versus 91.0% for the pooled human reference. Gaze detection remains near chance against 89.6% human accuracy. On actor disambiguation, reference-interface controls show that relational descriptions recover 5.55--13.25 points over static coordinates, confirming a substantial numeric-parsing penalty; yet visual boxes still lead every model by 1.15--6.50 points, exposing a residual unboxed actor-resolution gap. A binding-trap analysis shows models systematically select the wrong actor's action. ActionLens provides diagnostic measurements of these distinct failure modes across model families and scales for direct comparison. We release all data, code, and evaluation scripts at https://anonymous.4open.science/r/lmms-eval-2276