Accurate streaming detection of action changes in video sequences

Groupwise Selective State-Space Filtering for Accurate and Streaming Action Boundary Detection

Computer Vision and Pattern Recognition

Summary

Detecting where one action ends and another begins in long videos is important for analyzing activities without labeling what the actions are. The authors describe a method that looks at video features over time in groups to find these boundaries accurately. Their approach works without knowing the action types and can handle live video streams with little delay and limited memory. It was tested on different datasets and improved the detection of action transitions.

What this means in practice

  • For video analytics teams: Enable real-time segmentation of continuous video into meaningful intervals without needing action labels, improving downstream video analysis accuracy.
  • For security monitoring operators: Detect precise moments of activity changes in surveillance footage instantly to support quicker event response and investigation.

Authors

Mustafa Bora Çelik

Abstract

Action boundary detection partitions untrimmed video into intervals without assigning action classes. We present a boundary-detection adapter operating on pre-extracted video features, learning temporal representations via groupwise selective scans. Learned group fusion and temporal modeling convert these into transition scores, which are decoded into boundary timestamps. Trained with boundary-time supervision, the class-agnostic model is evaluated on Breakfast, GTEA, and 50Salads using temporal tolerances and bipartite matching, achieving boundary $F_1$ scores of 0.457, 0.622, and 0.611. A stateful variant enables feature-streaming inference with zero neural look-ahead, one-sample peak confirmation, and bounded memory. Downstream systems can subsequently assign s