Papers for

video content analysts

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Query-guided summarization improves locating events in long videos

Summarize Before Grounding: Query-Guided Chunk Condensation for Long-Video Temporal Grounding

Abstract: Video temporal grounding (VTG) aims to localize the video interval corresponding to a language query. Recent large vision-language models (LVLMs) show great potential in solving such a multi-modal reasoning task. However, long videos often contain large amounts of redundant information that disturbs LVLMs to mine query-relevant evidence. Instead of dense frame sampling which incurs prohibitive training memory, previous reinforcement learning with verifiable rewards (RLVR) works typically utilize sparse sampling, which makes training feasible but may miss critical evidence. In this paper, we propose a ``summarize before grounding'' framework (named ``SumGround'') for long-video temporal grounding. The key of SumGround is to perform query-guided chunk condensation to aggregate and retrieve query-relevant evidence. Specifically, we split the video into several chunks and perform two-level chunk condensation. First, we introduce query-guided latent summaries, which is represented as KV states of query-guided prompts, to compress redundant visual tokens into compact query-relevant chunk summaries. Furthermore, we design an associative summary retrieval scheme to rank and select chunk summaries that are most likely to contain the event interval. Both query-guided latent summary and associative summary retrieval schemes are enabled by RLVR. To reduce memory consumption, we propose a length-aware gradient gating module to selectively stop gradient back-propagated to visual tokens. Extensive experiments demonstrate that SumGround performs favorably against previous state-of-the-art methods across multiple downstream datasets, with remarkable gains on long videos.

Mon 28 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Finding the right moment in a long video that matches a spoken or written question is hard because videos have lots of extra, unimportant parts. The researchers created a new method that first summarizes parts of the video based on the question, then picks the most relevant summaries to find the exact moment. This reduces memory needs and helps the system focus better. Their method, called SumGround, works better than previous techniques, especially for long videos.
Open → 2609.34598v1

Livestream videos get smarter at matching products with show moments

Grounded Product Understanding in Livestream Videos

Abstract: E-commerce livestreams have emerged as an important channel for presenting products to online consumers, containing multiple products whose information is scattered in different moments. This poses significant challenges for downstream product understanding applications, such as product-centric livestream clipping, where models need to identify the product and its relevant segments for information gathering. However, existing benchmarks for general product understanding typically evaluate product retrieval and temporal localization in isolation, leaving the critical correspondence between product identity and temporal evidence largely unassessed. To address this limitation, we introduce GPUB, a large-scale benchmark comprising 3,000 livestream instances with quality-controlled multi-moment temporal annotations and a catalog of over 31K fashion products. GPUB supports three evaluation tasks: the main task Grounded Product Understanding (GPrU) requires jointly identifying the target product and localizing its supporting moments from a livestream video and a candidate product set; Product Retrieval and Product Moment Localization serve as two complementary subtasks. Evaluation of existing multimodal models shows that GPrU remains highly challenging, with the best-performing baseline achieving only 10.13% Pair mAP@.3. To narrow the performance gap, we further develop UniPro, a unified product understanding model that derives product-aligned and temporally structured representations from shared multimodal encoding, improving Pair mAP@.3 to 21.53% while achieving 37.23% Joint R@1@.3 on GPrU.

Thu 17 SeptComputer Vision and Pattern Recognition
The gist
Online shopping livestreams often show many products at different times, making it hard to tell which product relates to which video moment. The authors created GPUB, a big new dataset that connects fashion products to specific moments in livestreams to help computers learn this connection. They also built a new model called UniPro that does better at identifying products and the exact video parts that show them. Despite improvement, the task remains difficult for current technology.
Open → 2609.20508v1

Video sentiment analysis improves by splitting polarity and intensity tasks

Divide and Conquer: Mixture-of-Bottleneck Experts in Informative Ordinal Space for Video-based Multimodal Sentiment Analysis

Abstract: Video-based Multimodal sentiment analysis (MSA) must handle information from text, audio, and image sequence in human speaking videos, yet current methods often fail to integrate modalities with task awareness. Most models treat video sentiment prediction as a single task, overlooking its ordinal nature, and their fusion strategies struggle to capture diverse unique and synergic cues across modalities. To address these limitations, we adopt a divide-and-conquer perspective by reformulating MSA as an ordinal regression problem and decoupling it into polarity recognition and intensity prediction. Driven by information theory, we introduce a Mixture-of-Bottleneck (MoB) framework that assigns different latents to polarity- and intensity-specific experts for different modalities. With the learning of information bottleneck, each expert learns compact and task-relevant representations while filtering out redundancy and noise. A multimodal bottleneck routing fusion module then fuses these expert latents with hard mining strategy, guiding the prediction in the ordinal sentiment space. Extensive experiments on 4 MSA datasets and 4 language models show that MoB effectively leverages informative latents from diverse modalities and captures general sentiment structure. Beyond stronger performance, MoB comprehensively captures fine-grained intra- and inter-modal dynamics, enabling more trustworthy localization of nuanced video sentiment signals.

Wed 16 SeptMultimediaComputation and Language
The gist
Video sentiment analysis tries to understand feelings in videos using text, sound, and images together, but it is tricky because feelings have degrees and different parts work differently. The authors treat the problem as two tasks: recognizing positive or negative feelings and how strong they are. They introduce a method that uses small focused parts called bottlenecks to learn useful signals separately for these tasks from each kind of data. This helps the model better combine the different signals to predict feelings more accurately and understand subtle emotions in videos.
Open → 2609.18470v1