Livestream videos get smarter at matching products with show moments

Grounded Product Understanding in Livestream Videos

Computer Vision and Pattern Recognition

Summary

Online shopping livestreams often show many products at different times, making it hard to tell which product relates to which video moment. The authors created GPUB, a big new dataset that connects fashion products to specific moments in livestreams to help computers learn this connection. They also built a new model called UniPro that does better at identifying products and the exact video parts that show them. Despite improvement, the task remains difficult for current technology.

What this means in practice

  • For e-commerce platform developers: Automatically creating clips focused on specific products from livestreams to improve shopper experience.$Commercial implications: Enables building smarter livestream shopping tools that highlight products for customers, boosting sales and engagement.
  • For video content analysts: Improving tools that analyze livestreams by linking product mentions to the exact video segments where they appear.

Authors

Xinyu Zhang, Junjie Chen, Jiawei Ge, Qianlong Li, Libin Ma, Baokun Pan, Yahui Luo

Abstract

E-commerce livestreams have emerged as an important channel for presenting products to online consumers, containing multiple products whose information is scattered in different moments. This poses significant challenges for downstream product understanding applications, such as product-centric livestream clipping, where models need to identify the product and its relevant segments for information gathering. However, existing benchmarks for general product understanding typically evaluate product retrieval and temporal localization in isolation, leaving the critical correspondence between product identity and temporal evidence largely unassessed. To address this limitation, we introduce GPUB, a large-scale benchmark comprising 3,000 livestream instances with quality-controlled multi-moment temporal annotations and a catalog of over 31K fashion products. GPUB supports three evaluation tasks: the main task Grounded Product Understanding (GPrU) requires jointly identifying the target product and localizing its supporting moments from a livestream video and a candidate product set; Product Retrieval and Product Moment Localization serve as two complementary subtasks. Evaluation of existing multimodal models shows that GPrU remains highly challenging, with the best-performing baseline achieving only 10.13% Pair mAP@.3. To narrow the performance gap, we further develop UniPro, a unified product understanding model that derives product-aligned and temporally structured representations from shared multimodal encoding, improving Pair mAP@.3 to 21.53% while achieving 37.23% Joint R@1@.3 on GPrU.