Generating recommendation items in one step speeds up and improves accuracy
SPRINT: Single-Step Generative Recommendation via Average Probability Velocity
Information Retrieval
Summary
Recommendation systems usually predict what a user wants step-by-step, which takes time and computing power. This paper introduces a new way to generate an entire item recommendation in a single step using a concept called average probability velocity. The authors use a bidirectional Transformer model that predicts all parts of the item at once and then use a special training technique to keep the item coherent. Their experiments show this method is much faster and more accurate than previous step-by-step approaches.
What this means in practice
- •For recommender system engineers: Deploy single-step recommendation models that generate items faster and more accurately with fewer computation passes.
- •For online retail platform developers: Reduce recommendation latency by generating user-preferred items in one forward pass, improving user experience.
Authors
Zhuo Cai, Shoujin Wang, Peilin Zhou, Min Xu, Julian McAuley, Fang Chen
Abstract
Semantic ID (SID) based generative recommendation represents each item as a sequence of discrete tokens, and recommends by generating the SID of the item a user would like to interact with. Both dominant paradigms in this domain generally pay for generation token by token: autoregressive models decode the tokens left-to-right, while non-autoregressive models decode in parallel yet still need multiple rounds of refinement to stay competitive. Therefore, both generally spend multiple forward passes per item, a cost that is prohibitive in latency-sensitive recommender systems. We ask whether an item can be generated in a single forward pass, and answer it through a new perspective which we call average probability velocity. We view SID generation as a flow of token generation probabilities and characterize it by its average velocity over the whole generation process. We prove that this average velocity is fully determined by the average generation probability of each token. Therefore, we directly parameterize and learn the probabilities of all tokens in a single forward pass with a bidirectional Transformer. As these probabilities are generated independently across positions and the coherence among tokens is lost, we further design a dual-level flow contrastive objective to restore the coherence among an item's tokens. It contrasts the target SID against negative SIDs at both the token and SID levels. The token level ranks the generation probabilities of the target tokens above those of negative SIDs, while the SID level scores the tokens of each SID as a whole item for capturing token coherence of each item. Extensive experiments show that our model not only generates recommendations far more efficiently ($8.39-10.04\times$ speedup over the second-fastest AR/NAR method) but also attains superior recommendation accuracy ($7.77\%$ average improvement over the second-best.