EviDETR improves video search and highlight spotting from text queries

EviDETR: Preserving Query-Relevant Temporal Evidence for Moment Retrieval and Highlight Detection

Computer Vision and Pattern Recognition

Summary

Finding specific moments in videos and spotting highlights based on a text query is tricky because existing methods don’t always keep track of what parts of the video relate to the query. The authors propose EviDETR, a new system that better remembers relevant video segments during analysis and prediction. It uses special techniques to focus on important clips and combines information effectively to improve both search and highlight detection. EviDETR shows strong performance on several video datasets, making it easier to find and highlight moments in long videos.

What this means in practice

  • For video editing teams: Automatically identify and highlight relevant video segments from text descriptions to speed up content editing and review.
  • For multimedia content platforms: Enhance video search features by retrieving moments matching user queries and generating key highlights for improved browsing.$Commercial implications: Enables platforms to offer advanced video search and highlight features that attract users seeking personalized video experiences.

Authors

Haoran Sun, Yufan Li, Qichen Zhang, Haoran Zhao, Shuqi Wang

Abstract

Joint video moment retrieval and highlight detection requires identifying query-relevant temporal segments while estimating clip-level saliency, yet DETR-style pipelines do not explicitly preserve query-relevant evidence throughout encoding, decoding, and cross-task prediction. We propose EviDETR, an evidence-preserving framework with three components. Semantic-aware Feature Reweighting (SFR) enhances query-relevant clip representations through saliency estimation and cross-modal interaction. A Temporal Top-2 Mixture-of-Experts (TTop2MoE) decoder performs query-adaptive refinement via sparse expert routing. MR-to-HD (MR2HD) fusion transfers span-level retrieval evidence to clip-level highlight prediction through confidence-weighted multi-scale aggregation. Using CLIP+SlowFast features, EviDETR achieves 69.29 R1@0.5, 54.77 R1@0.7, and 48.41 Avg. mAP for moment retrieval on QVHighlights, together with 41.83 HD-mAP and 68.33 HIT@1. Strong results on TACoS and Charades-STA further demonstrate cross-dataset transferability.