Papers for

streaming service developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Improving generative recommender accuracy beyond initial predictions

Beyond the Beam: Constructive Repair and Candidate Completion for Generative Recommendation

Abstract: Generative recommenders retrieve items by generating identifiers, but a valid identifier can remain outside the beam after catalog expansion. This raises two connected questions: which failures can identifier assignment repair, and how should retrieval proceed beyond the initial beam? We characterize assignment repair with a fixed generator and retained old identifiers. Output-invariance certificates identify failures shared by all admissible assignments. Under a common effective prefix, coupled support and ranking constraints give the exact feasible interval of new-item counts for target recovery. Building on this characterization, Beyond the Beam (BB) obtains minimum-replacement repairs through an integral flow formulation, selects a shared map and adapts the generator. At inference, generative likelihood and collaborative evidence define one score for ranking, candidate priority and stopping. Retained prefix bounds guide candidate completion and certify its global Top-$K$ when the stopping condition is met. Exhaustive finite-catalog evaluation confirms construction in every feasible case. Across three Amazon Reviews categories and three random seeds, the full T5 procedure improves mean Recall@10 by 15.5--46.3% and NDCG@10 by 15.2--44.4% over the best-performing evaluated generative baseline for each dataset and metric. Matched controls show that shared construction and adaptation improve new-target ranking and certification efficiency on Beauty and Toys. Combined scoring and candidate completion improve NDCG@10 across all three datasets with both T5 and decoder-only LC-Rec.

Sun 27 SeptInformation Retrieval
The gist
Generative recommenders suggest items by creating their unique identifiers, but sometimes good choices are missed because they fall outside the first set of predictions. The authors study when and how these missed items can be recovered by adjusting assignments and generating more candidates beyond the usual limit. They propose a method called Beyond the Beam that fixes recommendations with minimal changes and uses combined signals to rank and decide when to stop searching. Tested on real Amazon data, their approach improves accuracy significantly compared to existing generative recommenders.
Open → 2609.33745v1

Audio description generation optimized for timing and content choices

What, When, and How: Audio Description as Constrained Global Optimization

Abstract: Audio Description (AD) makes movies accessible to blind and visually impaired audiences by narrating visual information in gaps between dialogue. Existing automatic AD systems largely treat generation as a local video-to-text problem, assuming that the content to describe and its temporal location are already provided. Realistic AD instead requires coupled decisions about what visual information is narratively important, when it can be spoken without interfering with dialogue, and how it should be formulated to fit within the available time. We formalize AD generation as a constrained optimization problem over these three decisions. Our hybrid system uses large language models to propose and ground visual elements, estimate their salience to the narrative, and generate compressed realizations. A mixed-integer linear program then jointly selects and schedules descriptions across a scene subject to temporal constraints. When evaluated on REFRAMED, a benchmark for realistic AD of movies, our approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new SOTA on narrative QA and temporally grounded metrics. Ablations show that explicit temporal constraints drive gains in placement, while salience estimation controls how much narratively useful content is retained. Improvements are concentrated on temporal and narrative measures rather than n-gram overlap, although a significant gap to professional describers remains.

Thu 24 SeptComputation and LanguageComputer Vision and Pattern Recognition
The gist
Making movies accessible to people who are blind means adding spoken descriptions of what’s happening on screen between dialogues. The authors show that this task isn’t just about describing what’s visible, but also about deciding what’s important to say, when to say it without overlapping dialogue, and how to say it briefly. They created a system that uses smart language models together with optimization techniques to pick and schedule descriptions for movie scenes. Their method improves on previous approaches in describing movies more thoughtfully and fitting descriptions better into the timing, although there’s still a gap compared to professional human describers.
Open → 2609.30121v1

GNN model adapts recommendations for individuals and groups

A Flexible Recommendation System for Individuals and Groups

Abstract: Group recommender systems typically rely on either aggregating individual preferences or treating groups as distinct meta-users. However, these methods often suffer from static aggregation strategies or data sparsity issues within group histories. This paper introduces a novel approach, that relies on a GNN-based architecture to learn a dual representation of each user's preferences, capturing their behavior as an independent individual from one side and as a member of a collective from the other side. By performing a differential analysis of these individual and group-oriented preferences, our system then determines the behavioral profile of each user when joining a group. Finally, specific preference aggregation strategies are defined to cope with the behavioral profiles of the users composing a group. Consequently, the system is equally capable of delivering precise recommendations to individuals and to arbitrary groups, effectively unifying the two traditional paradigms of recommendation. Experiments on synthetic data simulating diverse group settings and behaviors confirm the flexibility and relevance of the proposed approach compared to state-of-the-art methods.

Wed 23 SeptInformation Retrieval
The gist
Recommending items to groups is tricky because people’s tastes can change when they are together. The authors present a new method that learns each person’s preferences both alone and as part of a group. It compares these preferences to understand how people behave in groups and then uses this to offer better recommendations. Their system works well for both individuals and groups, solving some common problems with older methods.
Open → 2609.27998v1

Generative recommendation improves item suggestions using semantic and collaborative signals

Addressing Cross-Stage Decoupling of Semantic and Collaborative Signals in Generative Recommendation

Abstract: Generative recommendation reformulates sequential recommendation as autoregressive generation by encoding items into semantic tokens, enabling improved scaling capability and cross-domain generalization. However, existing generative recommender systems typically follow a two-stage pipeline, where item tokenization is largely dominated by textual semantics with limited incorporation of collaborative signals and interaction similarity, leading to code assignments that are misaligned with downstream generation. Conversely, the generation stage tends to overlook the original semantic information, as the code sequences are re-embedded based on interaction data. This cross-stage information decoupling limits semantic coherence and recommendation accuracy. To address this issue, we propose SCRec, a general framework that enhances cross-stage coherence through bidirectional information supplementation. Specifically, we introduce (i) collaborative-enhanced tokenization to explicitly inject textualized collaborative signals into semantic tokenization, without introducing additional alignment task, (ii) semantic-guided generation to dynamically recalibrate semantic priors with learnable code embeddings in generation stage, and (iii) manifold alignment to reconcile the geometric mismatch between the embedding space of discrete codebook indices and the dense continuous semantic space. These interrelated components form a general framework that aligns semantic and collaborative signals and enhances cross-stage information coherence, with minimal additional training and inference costs. Extensive experiments demonstrate the effectiveness, robustness, and generalizability of our proposed framework.

Sat 12 SeptInformation Retrieval
The gist
Recommendation systems help suggest things you might like based on your past choices and other users' behaviors. The authors found that current methods separate understanding an item's meaning from how people interact with it, making recommendations less accurate. They created a new approach called SCRec that better mixes item meanings with user behavior signals at all stages. This helps the system give smarter and more coherent recommendations without needing much extra training. Their tests show this method works well and can handle different recommendation scenarios.
Open → 2609.13678v1