Improving generative recommender accuracy beyond initial predictions
Beyond the Beam: Constructive Repair and Candidate Completion for Generative Recommendation
Information Retrieval
Summary
Generative recommenders suggest items by creating their unique identifiers, but sometimes good choices are missed because they fall outside the first set of predictions. The authors study when and how these missed items can be recovered by adjusting assignments and generating more candidates beyond the usual limit. They propose a method called Beyond the Beam that fixes recommendations with minimal changes and uses combined signals to rank and decide when to stop searching. Tested on real Amazon data, their approach improves accuracy significantly compared to existing generative recommenders.
What this means in practice
- •For ecommerce platform engineers: Enhance product recommendation systems to recover relevant items missed during initial prediction steps, improving customer choices.
- •For streaming service developers: Improve content recommendation by better handling new or expanded catalogs, leading to more accurate and diverse suggestions.
Authors
Zijun Zhao, Peng Zhang, Gang Zhang, Yuanchi Ma, Hui He, Zhendong Niu
Abstract
Generative recommenders retrieve items by generating identifiers, but a valid identifier can remain outside the beam after catalog expansion. This raises two connected questions: which failures can identifier assignment repair, and how should retrieval proceed beyond the initial beam? We characterize assignment repair with a fixed generator and retained old identifiers. Output-invariance certificates identify failures shared by all admissible assignments. Under a common effective prefix, coupled support and ranking constraints give the exact feasible interval of new-item counts for target recovery. Building on this characterization, Beyond the Beam (BB) obtains minimum-replacement repairs through an integral flow formulation, selects a shared map and adapts the generator. At inference, generative likelihood and collaborative evidence define one score for ranking, candidate priority and stopping. Retained prefix bounds guide candidate completion and certify its global Top-$K$ when the stopping condition is met. Exhaustive finite-catalog evaluation confirms construction in every feasible case. Across three Amazon Reviews categories and three random seeds, the full T5 procedure improves mean Recall@10 by 15.5--46.3% and NDCG@10 by 15.2--44.4% over the best-performing evaluated generative baseline for each dataset and metric. Matched controls show that shared construction and adaptation improve new-target ranking and certification efficiency on Beauty and Toys. Combined scoring and candidate completion improve NDCG@10 across all three datasets with both T5 and decoder-only LC-Rec.