Papers for

media analytics teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Counterfactual attention improves video event time marking accuracy

Counterfactual Attention Policy Distillation for Temporal Video Grounding

Abstract: Temporal video grounding is a key capability of advanced \emph{Multimodal Large Language Models} (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts in long videos. In this paper, we study this issue from the perspective of \emph{On-policy distillation} (OPD) and propose a new training regime for MLLMs termed \emph{Counterfactual Attention Policy Distillation} (CAPD). In particular, OPD is a viable solution for MLLMs via providing dense teacher supervision on student-generated trajectories. But its next-token based teacher-student distillation is hard to identify the specific video segments supporting each predicted timestamp, which is critical for temporal grounding. In this case, CAPD measures how masking each temporal group changes the teacher's output distribution. The resulting counterfactual influence calibrates the teacher's attention and weights token-level distillation, allowing the student to learn the temporal evidence that affects boundary prediction. To validate CAPD, we trained it on Qwen3-VL-8B-Instruct using only 2,500 samples for one epoch, and evaluated it on the TimeLens and multiple general video benchmarks. Experimental results show that CAPD improves average recall by 12.0\% relative to GRPO on TimeLens while preserving general video understanding, achieving comparable accuracy to the base model.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Finding exact moments when events happen in long videos is tricky because many actions look similar. The authors studied how to teach large language models to better identify those moments by focusing on the important video parts that influence their decisions. They created a new training method called CAPD, which looks at what happens if certain video segments are hidden to see how much they matter. This helped the model learn to pay attention to the right spots, improving its accuracy in timing events without losing its general video understanding skills.
Open → 2609.34581v1