Papers for

multimedia content platforms

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

EviDETR improves video search and highlight spotting from text queries

EviDETR: Preserving Query-Relevant Temporal Evidence for Moment Retrieval and Highlight Detection

Abstract: Joint video moment retrieval and highlight detection requires identifying query-relevant temporal segments while estimating clip-level saliency, yet DETR-style pipelines do not explicitly preserve query-relevant evidence throughout encoding, decoding, and cross-task prediction. We propose EviDETR, an evidence-preserving framework with three components. Semantic-aware Feature Reweighting (SFR) enhances query-relevant clip representations through saliency estimation and cross-modal interaction. A Temporal Top-2 Mixture-of-Experts (TTop2MoE) decoder performs query-adaptive refinement via sparse expert routing. MR-to-HD (MR2HD) fusion transfers span-level retrieval evidence to clip-level highlight prediction through confidence-weighted multi-scale aggregation. Using CLIP+SlowFast features, EviDETR achieves 69.29 R1@0.5, 54.77 R1@0.7, and 48.41 Avg. mAP for moment retrieval on QVHighlights, together with 41.83 HD-mAP and 68.33 HIT@1. Strong results on TACoS and Charades-STA further demonstrate cross-dataset transferability.

Fri 25 SeptComputer Vision and Pattern Recognition
The gist
Finding specific moments in videos and spotting highlights based on a text query is tricky because existing methods don’t always keep track of what parts of the video relate to the query. The authors propose EviDETR, a new system that better remembers relevant video segments during analysis and prediction. It uses special techniques to focus on important clips and combines information effectively to improve both search and highlight detection. EviDETR shows strong performance on several video datasets, making it easier to find and highlight moments in long videos.
Open → 2609.30724v1