Large scale dataset advances tracking of many animal species

VastMAT: A Large-Scale Multi-Category Benchmark for Multi-Animal Tracking

Computer Vision and Pattern Recognition

Summary

Tracking multiple animals at once helps scientists study how they move and interact, but most existing data focus on people or cars. The authors created VastMAT, a huge collection of videos showing many kinds of animals with detailed labels, to help computers get better at following animals in video. They tested current tracking methods and found it’s especially hard to track animals the system hasn’t seen before. They also added a simple technique that improves tracking by combining information about object centers and overlap.

What this means in practice

  • For wildlife monitoring teams: Track animal movements across many species in diverse environments using a large annotated video dataset to improve automated tracking systems.
  • For robotics developers: Improve robot vision systems to recognize and follow different animal types by training on a broad animal tracking dataset.

Authors

Zhizhen Li, Zan Wang, Huidong Peng, Bohan Tan, Shimin Shan, Yu Liu, Liang Peng

Abstract

Multi-animal tracking (MAT) supports the study of animal movement, behavior, and group interactions. However, general multi-object tracking (MOT) benchmarks primarily focus on pedestrians and vehicles, whereas dedicated MAT benchmarks remain limited in jointly supporting broad animal coverage, large-scale video data, and extensive within-video multi-instance association. To address this gap, we introduce VastMAT, which has four key characteristics: (1) Large scale. It comprises 2,947 videos with 1,002,562 annotated frames, totaling 27.85 hours. (2) Broad category coverage. These videos cover 337 animal categories with diverse morphologies and motion patterns. (3) Extensive instance annotations. It provides 3,663,248 bounding boxes and 22,883 identity trajectories---to our knowledge, the largest numbers of both among dedicated MAT benchmarks. (4) High-quality annotations. To ensure reliability, annotations undergo iterative expert review and correction, and quality is assessed through an independent reannotation audit. To systematically assess tracking performance and cross-category generalization, we establish Seen-category and category-disjoint Unseen-category protocols, and evaluate eight representative MOT methods under both protocols. Under these protocols, the highest baseline HOTA scores are 66.37\% and 52.90\%, respectively, highlighting the challenge of tracking unseen animals. To address the low-overlap association challenge revealed by our analysis, we propose Center-Distance-Augmented Association (CDA), a lightweight module that adaptively combines IoU with center similarity normalized by the boxes' own scales. Without additional training, CDA improves TrackTrack's HOTA by 1.58 and 1.31 percentage points under the two protocols, respectively. To facilitate further MAT research, we will publicly release our benchmark and code.