Video counting improves by learning when to update object states
When Should the Count Change? Learning State Maintenance for Causal Video Counting
Computer Vision and Pattern Recognition
Summary
Counting objects in videos is tricky because it's hard to tell if something new in the video is a new object, a returning one, or if an event has finished. The authors created a method called StaMina that keeps track of objects over time by learning when to update what it knows about each object. This method uses smart guessing about visibility and identity to count more accurately. They tested it on a large dataset and showed it counts better than earlier approaches.
What this means in practice
- •For video analytics teams: Improve continuous counting of objects or events in streams where distinguishing new and ongoing instances is crucial.
- •For surveillance system developers: Enhance automated counting features to better track persistent identities and completed events in security videos.
Authors
Pengyiang Liu, Dongyue Lyu, Junbo Niu, Zhongyue Shi, Jiahao Xie, Si Liu
Abstract
Continuous video counting requires distinguishing new observations from new objects or completed events. We introduce StaMina (State Maintenance), which learns to maintain counting state through state-conditioned updates. Recurrent visual context supports recognition; learned transitions maintain visibility, persistent identities, and completed-event records. A differentiable recurrence trains event transitions over legal paths constrained by count endpoints; visibility and association objectives train the object branch. A multi-source pipeline organizes 39.8K spatial queries and complementary event annotations into counting trajectories. On SVCBench, we evaluate counting adaptation with partial video overlap and held-out groups of linked annotations. Under prefix replay (Full) and persistent streaming (Stream), 4B and 8B models reach 41.9/36.4 and 44.9/38.2 Gaussian Precision Accuracy, respectively. The 8B model gains 10.9/3.2 points over Counting-SFT on the same queries. Matched-graph comparisons isolate phase conditioning and trajectory supervision, assessing training objectives alongside hard decisions. Online video benchmarks and count-conditioned decisions assess online understanding and task eligibility. Project Page: https://PLACEHOLDER.github.io/StaMina/