Embedding subspaces enable flexible multi-goal recommendations at scale
Embedding Subspace Partitioning for Dynamic Multi-Objective Retrieval
Information Retrieval
Summary
Recommender systems often need to balance different goals like showing relevant items and increasing business value, but traditional methods mix all goals into one fixed understanding. The authors propose splitting the understanding into separate parts, each focused on a specific goal, so the system can adjust priorities on the fly without retraining. Their method works efficiently on large-scale systems by doing one search over combined data rather than multiple searches. Tests show this approach better balances goals and improved user engagement in a large job matching platform.
What this means in practice
- •For recommendation engineers: Enable dynamic adjustment of recommendation priorities without retraining by separating embeddings into task-specific parts.
- •For search infrastructure teams: Implement efficient multi-objective retrieval using a single combined search index to avoid costly multiple nearest neighbor systems.
- •For advertising platform developers: Adjust trade-offs between user relevance and revenue goals dynamically in ad recommendation systems for better performance under shifting priorities.$Commercial implications: Allows development of configurable ad recommendation products that adapt business goals without retraining, improving monetization flexibility.
Authors
Shaobo Zhang, Alice Leung, Yunxiang Ren, Ping Liu, Yuchin Juan, Qianqi Shen, Benjamin Le, Jianqiang Shen, Chengming Jiang, Ko-Cheng Wang, Vidya Krishnamurthy, Caleb Johnson, Fedor Borisyuk, Luke Simon, Jingwei Wu, Wenjing Zhang
Abstract
Modern industrial recommender systems must optimize across competing objectives, balancing semantic relevance with business metrics such as engagement and revenue. While bi-encoders dominate large-scale retrieval due to their efficiency, they collapse these heterogeneous signals into a single static embedding space. This design creates a fundamental limitation: once trained, the retriever cannot adapt to shifting objective priorities at serving time without retraining. Moreover, joint optimization with multi-objective losses often induces interference between objectives, leading to suboptimal trade-offs. We propose Embedding Subspace Partitioning (ESP), a retrieval framework that decomposes the embedding into task-aware subspaces and replaces the single dot product with a weighted sum of per-subspace similarities, whose weights are tunable at serving time. For Transformer bi-encoders, ESP uses the model's native end-of-sequence token as a segment delimiter, with segment-aware attention masking and position encoding resets to guarantee subspace isolation in a single forward pass. Serving is performed via GPU-accelerated exhaustive kNN over one concatenated index, eliminating the need for per-objective Approximate Nearest Neighbor (ANN) infrastructure required by multi-head approaches. We evaluate ESP on an open-source benchmark built from MS MARCO. A single ESP model traces a broad Pareto frontier, consistently outperforming strong multi-task baselines across diverse operating points. In LinkedIn's job matching platform (70M+ weekly users), ESP enabled dynamic retrieval reconfiguration and delivered significant key business metric lifts.