Multi-camera pedestrian location without calibration or training needed

CAT-Free: Multi-View Pedestrian Localization without Calibration, Annotations, or Target-Scene Training via Adaptive Geometric Filtering

Computer Vision and Pattern Recognition

Summary

Tracking where people are in big spaces using many cameras usually requires setting up each camera carefully, marking exact positions, or training the system beforehand. The authors developed a method called CAT-Free that only needs synchronized video from cameras and no extra setup or training. It figures out camera positions automatically from the video and combines information to locate people. Because the automatic setup can be inaccurate, it uses special filters that remove unreliable location guesses, making the system work well without prior calibration.

What this means in practice

  • For video surveillance teams: Deploy pedestrian localization in new environments without manual camera setup or pre-existing annotations by using synchronized camera videos only.
  • For smart building operators: Enable wide-area indoor pedestrian monitoring without costly calibration or training when adding or moving cameras in a building.

Authors

Taigo Sakai, Hiroki Kouno, Naoki Kato, Kazuhiro Hotta

Abstract

Multi-camera pedestrian localization is useful for wide-area monitoring in public and commercial spaces. However, deploying these systems often requires considerable setup for each new environment. Existing methods typically require camera calibration, position annotations, or target-scene training. CAT-Free removes all three requirements. It uses synchronized RGB video as its only scene-specific input. Camera configuration is estimated directly from the video. Pedestrian locations are then estimated by combining observations from multiple cameras. Automatic camera estimation is not always accurate. This can produce unreliable pedestrian locations. CAT-Free therefore introduces two adaptive geometric filters. They remove unreliable position estimates. Their thresholds are estimated from each input sequence. CAT-Free achieves 82.5, 84.5, and 65.7 MODA on WildTrack, MultiviewX, and GMVD. It uses no supplied calibration, position annotations, or target-scene training. Published methods using such scene-specific information report 88.2--95.0 MODA on WildTrack and 83.9--96.5 on MultiviewX under their respective protocols. CAT-Free also transfers without retuning. It reaches 74.9 MODA on four additional sequences and 78.6 on an unseen 8-camera installation. Finally, localization uncertainty predicts MODA with $r=-0.98$. This provides a label-free estimate of localization reliability.