Generative recommendation improves item suggestions using semantic and collaborative signals
Addressing Cross-Stage Decoupling of Semantic and Collaborative Signals in Generative Recommendation
Information Retrieval
Summary
Recommendation systems help suggest things you might like based on your past choices and other users' behaviors. The authors found that current methods separate understanding an item's meaning from how people interact with it, making recommendations less accurate. They created a new approach called SCRec that better mixes item meanings with user behavior signals at all stages. This helps the system give smarter and more coherent recommendations without needing much extra training. Their tests show this method works well and can handle different recommendation scenarios.
What this means in practice
- •For ecommerce platform teams: Improve product recommendation quality by integrating user interaction data with item semantics throughout the recommendation process.
- •For streaming service developers: Enhance personalized content suggestions by aligning item descriptions with collaborative user preferences in generative recommendation systems.
Authors
Jiayi Dan
Abstract
Generative recommendation reformulates sequential recommendation as autoregressive generation by encoding items into semantic tokens, enabling improved scaling capability and cross-domain generalization. However, existing generative recommender systems typically follow a two-stage pipeline, where item tokenization is largely dominated by textual semantics with limited incorporation of collaborative signals and interaction similarity, leading to code assignments that are misaligned with downstream generation. Conversely, the generation stage tends to overlook the original semantic information, as the code sequences are re-embedded based on interaction data. This cross-stage information decoupling limits semantic coherence and recommendation accuracy. To address this issue, we propose SCRec, a general framework that enhances cross-stage coherence through bidirectional information supplementation. Specifically, we introduce (i) collaborative-enhanced tokenization to explicitly inject textualized collaborative signals into semantic tokenization, without introducing additional alignment task, (ii) semantic-guided generation to dynamically recalibrate semantic priors with learnable code embeddings in generation stage, and (iii) manifold alignment to reconcile the geometric mismatch between the embedding space of discrete codebook indices and the dense continuous semantic space. These interrelated components form a general framework that aligns semantic and collaborative signals and enhances cross-stage information coherence, with minimal additional training and inference costs. Extensive experiments demonstrate the effectiveness, robustness, and generalizability of our proposed framework.