Papers for

recommendation engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Embedding subspaces enable flexible multi-goal recommendations at scale

Embedding Subspace Partitioning for Dynamic Multi-Objective Retrieval

Abstract: Modern industrial recommender systems must optimize across competing objectives, balancing semantic relevance with business metrics such as engagement and revenue. While bi-encoders dominate large-scale retrieval due to their efficiency, they collapse these heterogeneous signals into a single static embedding space. This design creates a fundamental limitation: once trained, the retriever cannot adapt to shifting objective priorities at serving time without retraining. Moreover, joint optimization with multi-objective losses often induces interference between objectives, leading to suboptimal trade-offs. We propose Embedding Subspace Partitioning (ESP), a retrieval framework that decomposes the embedding into task-aware subspaces and replaces the single dot product with a weighted sum of per-subspace similarities, whose weights are tunable at serving time. For Transformer bi-encoders, ESP uses the model's native end-of-sequence token as a segment delimiter, with segment-aware attention masking and position encoding resets to guarantee subspace isolation in a single forward pass. Serving is performed via GPU-accelerated exhaustive kNN over one concatenated index, eliminating the need for per-objective Approximate Nearest Neighbor (ANN) infrastructure required by multi-head approaches. We evaluate ESP on an open-source benchmark built from MS MARCO. A single ESP model traces a broad Pareto frontier, consistently outperforming strong multi-task baselines across diverse operating points. In LinkedIn's job matching platform (70M+ weekly users), ESP enabled dynamic retrieval reconfiguration and delivered significant key business metric lifts.

Thu 24 SeptInformation Retrieval
The gist
Recommender systems often need to balance different goals like showing relevant items and increasing business value, but traditional methods mix all goals into one fixed understanding. The authors propose splitting the understanding into separate parts, each focused on a specific goal, so the system can adjust priorities on the fly without retraining. Their method works efficiently on large-scale systems by doing one search over combined data rather than multiple searches. Tests show this approach better balances goals and improved user engagement in a large job matching platform.
Open → 2609.30601v1

Multi task learning improves recommendation accuracy and satisfaction

Learned Cross-Task Relationships in Multi-Task Models

Abstract: We propose a framework that learns cross-task relationships in multi-task models by approximating the joint distribution of task labels through targeted pairwise relationships. This approach improves performance via transfer learning and enhances information extraction without the intractable complexity of modeling the full joint space. Although our framework applies to any multi-task system, we demonstrate its efficacy within YouTube's production recommendation systems. Experiments across the Notifications, Homepage, and Watch Next surfaces show improvements in both accuracy and user satisfaction metrics. Finally, we propose a workflow template to facilitate broader future implementation.

Wed 23 SeptArtificial Intelligence
The gist
Many online services try to do several related tasks at once, like recommending videos and sending notifications. The authors introduce a way to help computers understand how these tasks relate to each other by looking at pairs of tasks rather than everything at once. This approach makes the system better at learning and sharing useful information, improving recommendation performance. They tested their method on YouTube’s recommendations and found it made users happier with the suggestions they received.
Open → 2609.28776v1

Dataset recommender system helps choose evaluation data for recommender algorithms

FINALLY: A Dataset Recommender System for Recommender-Systems Research

Abstract: Dataset selection shapes the empirical conditions under which recommender-system algorithms are evaluated, yet existing tools provide limited support for constructing complete dataset sets that jointly satisfy experimental constraints and set-level selection objectives. To address this problem, I developed FINALLY, a web-based dataset recommender for constructing configurable dataset sets for offline recommender-systems evaluations. FINALLY combines required datasets, candidate-pool restrictions, metadata filters, configurable target-set sizes, Random selection, and diverse and non-diverse strategies based on adapted Effective Covariance and Convex Hull objectives. I evaluated FINALLY through 420 recommendation runs across ten systematically varied configurations. All evaluated dataset sets satisfied the applicable target-size, duplicate-avoidance, snapshot-membership, required-dataset, and metadata-filter requirements. All 40 deterministic strategy--configuration combinations were reproducible. Both the Effective-Covariance-based and Convex-Hull-based strategies produced the expected diverse-versus-non-diverse score ordering in all ten configurations. Under their corresponding objectives, the diverse strategies produced scores above all 30 configuration-specific Random results, whereas the non-diverse strategies produced scores below all 30 Random results. These results establish technical consistency for the evaluated FINALLY workflow and show that the implemented strategies follow their intended optimization directions within the investigated configuration space. They do not establish the scientific suitability, global optimality, or practical superiority of the generated selections.

Tue 8 SeptInformation Retrieval
The gist
Choosing the right datasets to test recommendation algorithms is important but challenging. The authors created FINALLY, a web tool that helps pick groups of datasets meeting specific rules and goals. They tested it thoroughly and found the tool consistently selects datasets as intended, including diverse and non-diverse groups. However, they note this does not prove which selections are best for all uses. The tool helps system builders set up fair and varied testing scenarios for recommendation algorithms.
Open → 2609.08941v1