Vision language models struggle to link actions correctly over time

ActionLens: Diagnosing Spatial-Temporal Binding Failures in Vision-Language Models

Computer Vision and Pattern RecognitionArtificial IntelligenceComputation and Language

Summary

Many current video understanding AI systems have trouble correctly matching who is doing what and when in videos. The authors created ActionLens, a large set of detailed multiple-choice questions designed to test specific problems in this area, like recognizing different people’s actions and where they are looking. They found that even the best models perform much worse than humans, especially when identifying gaze or matching actions to the right person. The benchmark helps reveal exactly where these models fail so future improvements can be better targeted.

What this means in practice

Authors

Gueter Josmy Faure, Min-Hung Chen, Hao Ping Wang, Timothée Lardy, Hung-Ting Su, Winston H. Hsu

Abstract

Video-capable vision-language models score above 80\% on popular benchmarks yet struggle with spatial-temporal binding: associating the right action with the right person at the right moment. We introduce ActionLens, a diagnostic benchmark of 6,701 multiple-choice video questions spanning five targeted diagnostics: transition detection, actor-specific identification, concurrent action binding, directed interaction reasoning, and gaze detection. Ground-truth answers are derived deterministically from 1.58 million per-second, per-person annotations. Fourteen rounds of human quality engineering raised answer clarity from 53% to above 90% human accuracy. Across 20 VLMs, the full-set leader scores 68.8%; on the human-reviewed subset, it scores 65.9% versus 91.0% for the pooled human reference. Gaze detection remains near chance against 89.6% human accuracy. On actor disambiguation, reference-interface controls show that relational descriptions recover 5.55--13.25 points over static coordinates, confirming a substantial numeric-parsing penalty; yet visual boxes still lead every model by 1.15--6.50 points, exposing a residual unboxed actor-resolution gap. A binding-trap analysis shows models systematically select the wrong actor's action. ActionLens provides diagnostic measurements of these distinct failure modes across model families and scales for direct comparison. We release all data, code, and evaluation scripts at https://anonymous.4open.science/r/lmms-eval-2276