Papers for

streaming service engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

SPADE metric measures truly surprising recommendations beyond popularity and similarity

SPADE: Escaping the Popularity-Similarity Frontier to Measure Serendipitous Recommendations

Abstract: Recommender systems engineer serendipity to foster active exploration and break predictable consumption cycles. The problem with existing offline beyond-accuracy metrics is that they often either isolate historical similarity or global popularity. We aim to design an evaluation metric that examines similarity, popularity, and actual user relevance. To achieve this, we introduce SPADE (Serendipitous Pareto Distance Evaluation). SPADE maps all items into a two-dimensional space to directly calculate a user-specific Pareto frontier of maximally popular and historically similar items. The final serendipity score is then computed by averaging the minimum Euclidean distance from this boundary strictly for the correctly recommended test-set items. Evaluating SPADE across five datasets and five baseline algorithms confirms its effectiveness; our results show that the metric successfully prevents algorithms from exploiting beyond-accuracy measures with irrelevant or non-personalized recommendations, reliably isolating serendipitous discoveries.

Fri 25 SeptInformation RetrievalArtificial IntelligenceMachine Learning
The gist
Recommender systems often suggest items based on popularity or similarity, which can make recommendations predictable. The authors created SPADE, a new way to evaluate how surprising or serendipitous recommendations are by considering popularity, similarity, and what the user actually likes. SPADE works by mapping items on a chart and measuring how far recommended items are from popular or similar items that a user already knows. Testing this across several datasets showed SPADE reliably identifies recommendations that are both relevant and genuinely unexpected.
Open → 2609.31164v1

Latent semantic fusion improves sequential recommendation accuracy

LSF-SR: Latent Semantic Fusion for Sequential Recommendation via Flow-based Conditional Variational Autoencoders

Abstract: Sequential recommendation aims to predict users' future interests from their historical interactions. Although Large Language Models (LLMs) capture rich item semantics, existing methods often struggle to align collaborative signals with textual semantic knowledge. As a result, the learned item representations fail to capture the complementary strengths of both signals, leading to suboptimal recommendation quality. To address this limitation, we propose Latent Semantic Fusion for Sequential Recommendation via Flow-based Conditional Variational Autoencoders (LSF-SR), a novel framework that uses a Conditional Variational Autoencoder (CVAE) with Normalizing Flows to fuse item ID embeddings and LLM-generated semantic signals. At the core of LSF-SR is a conditional fusion module augmented with planar or radial flows. This module learns a flexible latent space that encourages items with similar semantic profiles to cluster together within the latent manifold. Through extensive experiments on five public benchmark datasets, we demonstrate that LSF-SR consistently outperforms state-of-the-art baselines, achieving gains of up to 12.98% and 14.13% in Recall@20 and NDCG@20, respectively.

Thu 24 SeptInformation Retrieval
The gist
Sequential recommendation systems try to guess what items you might like next based on what you liked before. The authors found that combining different types of information about items, like their unique IDs and the meanings captured by text-based language models, can be tricky. They developed a new method called LSF-SR that uses a special kind of machine learning model to blend these data types more effectively. This approach helps group similar items together in a better way, which makes the recommendations more accurate.
Open → 2609.29815v1

Semantic user profiling methods compared for streaming recommendations

When LLM-Based User Profiling Adds Value in Production Streaming Recommendation

Abstract: Personalized recommendation depends critically on how user representations are constructed from historical behavior. Two paradigms have emerged for constructing semantic user profiles in content-based recommendation. First, aggregate methods derive user representations as numerical aggregates of semantic item embeddings. Second, LLM-based methods generate natural-language summaries of user preferences and encode them through a text encoder. Each paradigm can be combined with temporal disentanglement of recent versus historical behavior. LLM-based profile generation is significantly more expensive than aggregate approaches, raising the question of when this additional cost is justified. We present a systematic comparison of four semantic user-profiling strategies, factorially crossed across representation type and temporal handling, evaluated on a real-world production dataset. The comparison reveals how these strategies differ across user behavior types, across both accuracy and beyond-accuracy dimensions of recommendation quality, and across the temporal-window setting that governs the disentanglement.

Wed 23 SeptInformation Retrieval
The gist
Online services often suggest content based on what users liked before, but creating user profiles from past actions varies in complexity and cost. The authors compare two main ways to build these profiles: one that averages numerical representations of items, and another that uses large language models (LLMs) to write natural descriptions of user interests. They tested these methods on a real streaming service dataset, looking at how well each works depending on the user’s behavior and the timing of actions. Their findings help decide when the expensive LLM-based profiling is worth using compared to simpler methods.
Open → 2609.27183v1

Physics based system forecasts LEO satellite internet quality anywhere

Reading the Sky to Forecast the Ground: Physics-Informed Link-State Forecasting for LEO Networks at Any Location

Abstract: In this paper, we introduce Gnomon, a physics-informed system that forecasts user-perceived low-Earth-orbit (LEO) downlink throughput, uplink throughput, and round-trip time (RTT) under different levels of trace availability. Gnomon's physics layer reconstructs the serving geometry and four-leg bent-pipe attenuation from public weather, orbital, routing, and licensing data. Based on what is available, Gnomon conditions on the target terminal's own history (Mode 1), measurements from nearby publicly reachable dishes (Mode 2), or the physical covariates alone (Mode 3) to predict the link state: Modes 1 and 2 share a fine-tuned time-series foundation model, while Mode 3 uses a compact boosted-tree estimator. All three modes expose a common output interface and can be selected without retraining. We evaluate Gnomon using 8,260 minutes of 1 Hz measurements collected at nine sites across five states in the U.S. We train on three sites and hold out the remaining six sites and their serving beams. On these unseen sites, the own-trace mode reduces downlink-throughput and RTT prediction error by 17% and 11% relative to the strongest published baseline and, to our knowledge, provides the first LEO uplink forecasts. The neighbor-trace mode requires no on-site hardware, while the covariate-only mode reduces downlink-throughput and RTT error by 24.6% and 78.8% relative to the only prior covariate-only forecaster. Moreover, experiments show that Gnomon provides calibrated quantile bands and improves adaptive-bitrate streaming driven over real TCP flows on replayed Starlink links.

Tue 22 SeptNetworking and Internet Architecture
The gist
Low-Earth-orbit (LEO) satellites provide internet, but their connection quality can vary with weather and location, making it hard to predict. This paper presents Gnomon, a system that uses physics and public data sources to forecast download speed, upload speed, and latency for satellite internet users, even with limited local measurements. Gnomon offers three different ways to predict network quality based on what data is available, improving accuracy on new or unseen locations. It also helps streaming video perform better by adjusting quality based on these forecasts.
Open → 2609.26696v1

Feature decorrelation improves balance in sequential item recommendations

A Redundancy Reduction Approach for Controllable Sequential Recommendations

Abstract: Sequential recommendation must operate under long-tailed item distributions and popularity-driven concentration, often forcing practitioners to trade short-list accuracy against long-tail exposure. In this work, we study feature decorrelation as a mechanism for shaping representation geometry in dot-product sequential recommenders, and analyze how this, in turn, affects popularity-driven concentration. We propose a decorrelation-regularized training framework that augments next-item prediction with an auxiliary redundancy-reduction term, and instantiate it with BT-SR, which uses the Barlow Twins objective. To form label-consistent positive pairs without synthetic corruptions, we pair user histories that share the same next-item target. Beyond accuracy, we provide a geometric analysis showing how decorrelation suppresses shared low-rank directions in the user representation space that can give popular items a global scoring advantage, and we introduce a bucket-based alignment concentration metric to quantify this effect. Experiments on five public benchmarks show that BT-SR consistently improves next-item ranking quality, while the decorrelation strength acts as a simple control knob that reallocates accuracy across head and tail items, enabling accuracy-exposure trade-offs. Our analysis also reveals that the impact on head-vs-tail exposure differs across datasets, reflecting interactions between decorrelation and data temporal structure.

Sun 20 SeptInformation Retrieval
The gist
Items people see recommended often favor popular choices, making it hard for less-known items to appear. The authors studied a new way to reduce overlap in recommendation features, which helps spread attention more evenly across popular and less popular items. They created a training method that encourages this feature variety by comparing user histories with similar next-item interests. This method improved recommendation rankings and allows control over how much focus is given to popular versus less popular items.
Open → 2609.23849v1