Vision transformers match risky tackle detection recall in football videos
Revisiting Risky Tackle Detection with Vision Transformers
Computer Vision and Pattern Recognition
Summary
Detecting risky tackles in football practice videos can help improve player safety. The authors revisited previous work that used a video-based vision transformer model to spot risky tackles and confirmed the original reported accuracy numbers by reproducing the method exactly. They found that adjusting brightness was the most important factor in the model's success, while other changes like rotation or flipping were less helpful. Without these adjustments, the model performed worse than simpler methods previously used.
What this means in practice
- •For sports analytics teams: Use vision transformer models with brightness augmentation to detect risky tackles in football practice videos for injury prevention analysis.
- •For video surveillance engineers: Implement video classification pipelines that emphasize brightness adjustments to improve detection of specific events in sports recordings.
Authors
Syed Ahsan Masud Zaidi, Lior Shamir, Scott Dietrich
Abstract
This paper is a Track 2 reproducibility companion to an ICPR 2026 study on risky tackle detection in American football prac- tice videos. The original work fine-tuned a Video Vision Transformer (ViViT) on 733 clips labeled with the SATT-3 rubric. It used focal loss, Taguchi L18 augmentation, and 5-fold cross-validation. It reported risky- class recall of 0.67 and risky-class F1 of 0.59. This companion documents the released artifact and traces those numbers to specific scripts, fold out- puts, and aggregation files. The reproduced headline is run_15. It com- bines Gaussian noise with static brightness decrease and uses no rotation and no flip. Its fold-mean risky recall is 0.667 and its fold-mean risky F1 is 0.588. These values match the published headline after rounding. The ablation shows that brightness is the dominant factor. Its risky-recall main-effect range is 0.055, which is larger than the ranges for rotation, flip, and noise. Without augmentation, ViViT reaches risky recall of 0.545 and does not exceed the C3D baseline of 0.583. The raw clips show iden- tifiable student athletes, so they cannot be redistributed. The artifact provides a public sample for pipeline checks and a controlled route for full-data review.