Counterfactual attention improves video event time marking accuracy

Counterfactual Attention Policy Distillation for Temporal Video Grounding

Computer Vision and Pattern Recognition

Summary

Finding exact moments when events happen in long videos is tricky because many actions look similar. The authors studied how to teach large language models to better identify those moments by focusing on the important video parts that influence their decisions. They created a new training method called CAPD, which looks at what happens if certain video segments are hidden to see how much they matter. This helped the model learn to pay attention to the right spots, improving its accuracy in timing events without losing its general video understanding skills.

What this means in practice

  • For video platform developers: Enhance automated tools that mark event start and end times to improve content indexing and retrieval accuracy in video libraries.
  • For media analytics teams: Improve temporal analysis of video data for detailed event detection in long videos, benefiting content summarization and highlight generation.

Authors

Shaobo Ju, Haiyang Yu, Xuecheng Wu, Qiong Wu, Jiacong Wang, Fan Shi, Peng Jun, Yiyi Zhou

Abstract

Temporal video grounding is a key capability of advanced \emph{Multimodal Large Language Models} (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts in long videos. In this paper, we study this issue from the perspective of \emph{On-policy distillation} (OPD) and propose a new training regime for MLLMs termed \emph{Counterfactual Attention Policy Distillation} (CAPD). In particular, OPD is a viable solution for MLLMs via providing dense teacher supervision on student-generated trajectories. But its next-token based teacher-student distillation is hard to identify the specific video segments supporting each predicted timestamp, which is critical for temporal grounding. In this case, CAPD measures how masking each temporal group changes the teacher's output distribution. The resulting counterfactual influence calibrates the teacher's attention and weights token-level distillation, allowing the student to learn the temporal evidence that affects boundary prediction. To validate CAPD, we trained it on Qwen3-VL-8B-Instruct using only 2,500 samples for one epoch, and evaluated it on the TimeLens and multiple general video benchmarks. Experimental results show that CAPD improves average recall by 12.0\% relative to GRPO on TimeLens while preserving general video understanding, achieving comparable accuracy to the base model.