Papers for

video platform engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Semantic matching improves search and summaries for coding videos

Intelligent Semantic Matching (ISM) for Video Tutorial Search using Transformer Models

Abstract: The rise in the number and diversity of available software development video tutorials has enhanced digital learning for developers but also introduced challenges in locating relevant content efficiently. Existing video search methods, including keyword-based approaches and tools like CodeTube and TechTube, rely primarily on retrieval algorithms such as BM25, which fail to capture the semantic nuances and user intentions behind search queries. To address these limitations, we introduce ISM, an approach that uses SBERT to generate semantically rich vectors from video tutorial transcripts to improve the search for programming video tutorials. By segmenting transcripts and implementing a re-ranking process, ISM effectively preserves context and enhances the relevance of search results. Additionally, ISM generates informative video summaries using GPT-4, allowing developers to quickly assess the relevance of video content. To evaluate our approach, we first performed a quantitative study comparing ISM with the baseline TechTube. The results revealed that ISM performs better in both video retrieval and fragment identification, achieving a Hit@5 score of 0.95 and an average F1 score of 0.70 compared to the baseline's 0.58 and 0.52, respectively. We also performed a user study, which revealed that users strongly preferred the semantic matching capabilities and AI-generated summaries of our approach. This work advances the state-of-the-art in programming video tutorial search and summarization by offering more nuanced and user-aligned retrieval and summarization mechanisms.

Fri 11 SeptSoftware Engineering
The gist
Finding the right software tutorial videos can be hard because usual searches only look for keywords, missing the true meaning behind questions. The authors created ISM, which understands the meaning of video transcripts to better match search queries. It breaks videos into parts and reorders results to keep the context clear. ISM also uses AI to make helpful video summaries so developers can quickly decide if a video is useful. Their tests showed ISM finds better matches and users liked the smart search and summaries.
Open 2609.12921v1

Foundation models improve video search by adapting processing depth

Beyond Similarity: Foundation Models as an Efficient Backbone for Training-Free Composed Video Retrieval

Abstract: Composed video retrieval (CoVR) searches a gallery for the target video that realizes a natural-language modification of a source clip. However, at gallery scale, this creates a fundamental tension: compact embeddings enable efficient, reusable search but can miss the transient actions, state changes, and subtle constraints that demand fine-grained video reasoning, whereas applying large multimodal models uniformly sacrifices scalability. To address these limitations, we propose that frozen foundation models should instead occupy complementary roles, with inference depth adapted to query difficulty. Based on this premise, we introduce \methodname{}, a framework for training-free \methodexpansion{}. Specifically, a composed-query embedding first searches reusable video-only gallery representations; uncertain queries undergo bounded reranking and candidate expansion; ambiguous edits trigger target-description generation; and only close leading candidates reach multimodal verification. To support these roles, frame selection, spatial resolution, and time cues are adapted to each stage. Across complete target-gallery evaluations, our method reaches state-of-the-art performance among training-free approaches, with 89.55 and 93.43 R@1 on Dense-WebVid-CoVR and CoVR-R, respectively (with more than +35\% and +25\% absolute margins to the closest counterpart). These results show that adaptively orchestrating foundation-model capabilities can combine scalable retrieval with fine-grained reasoning without task-specific training. The source code and all relevant guidelines are available on https://github.com/demidovd98/CoVRAGE.

Wed 9 SeptComputer Vision and Pattern Recognition
The gist
Finding a video based on a change you describe from another video is hard because videos are complex and big. The authors show that using large pre-trained AI models in smart steps—starting simple and getting more detailed only when needed—makes searching much faster and more precise. Their method doesn’t require extra training and works well on big video libraries. This helps computers find videos that match complex descriptions without slowing down too much.
Open 2609.10008v1