Papers for
streaming service engineers
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
SPADE metric measures truly surprising recommendations beyond popularity and similarity
SPADE: Escaping the Popularity-Similarity Frontier to Measure Serendipitous Recommendations
Abstract: Recommender systems engineer serendipity to foster active exploration and break predictable consumption cycles. The problem with existing offline beyond-accuracy metrics is that they often either isolate historical similarity or global popularity. We aim to design an evaluation metric that examines similarity, popularity, and actual user relevance. To achieve this, we introduce SPADE (Serendipitous Pareto Distance Evaluation). SPADE maps all items into a two-dimensional space to directly calculate a user-specific Pareto frontier of maximally popular and historically similar items. The final serendipity score is then computed by averaging the minimum Euclidean distance from this boundary strictly for the correctly recommended test-set items. Evaluating SPADE across five datasets and five baseline algorithms confirms its effectiveness; our results show that the metric successfully prevents algorithms from exploiting beyond-accuracy measures with irrelevant or non-personalized recommendations, reliably isolating serendipitous discoveries.
Latent semantic fusion improves sequential recommendation accuracy
LSF-SR: Latent Semantic Fusion for Sequential Recommendation via Flow-based Conditional Variational Autoencoders
Abstract: Sequential recommendation aims to predict users' future interests from their historical interactions. Although Large Language Models (LLMs) capture rich item semantics, existing methods often struggle to align collaborative signals with textual semantic knowledge. As a result, the learned item representations fail to capture the complementary strengths of both signals, leading to suboptimal recommendation quality. To address this limitation, we propose Latent Semantic Fusion for Sequential Recommendation via Flow-based Conditional Variational Autoencoders (LSF-SR), a novel framework that uses a Conditional Variational Autoencoder (CVAE) with Normalizing Flows to fuse item ID embeddings and LLM-generated semantic signals. At the core of LSF-SR is a conditional fusion module augmented with planar or radial flows. This module learns a flexible latent space that encourages items with similar semantic profiles to cluster together within the latent manifold. Through extensive experiments on five public benchmark datasets, we demonstrate that LSF-SR consistently outperforms state-of-the-art baselines, achieving gains of up to 12.98% and 14.13% in Recall@20 and NDCG@20, respectively.
Semantic user profiling methods compared for streaming recommendations
When LLM-Based User Profiling Adds Value in Production Streaming Recommendation
Abstract: Personalized recommendation depends critically on how user representations are constructed from historical behavior. Two paradigms have emerged for constructing semantic user profiles in content-based recommendation. First, aggregate methods derive user representations as numerical aggregates of semantic item embeddings. Second, LLM-based methods generate natural-language summaries of user preferences and encode them through a text encoder. Each paradigm can be combined with temporal disentanglement of recent versus historical behavior. LLM-based profile generation is significantly more expensive than aggregate approaches, raising the question of when this additional cost is justified. We present a systematic comparison of four semantic user-profiling strategies, factorially crossed across representation type and temporal handling, evaluated on a real-world production dataset. The comparison reveals how these strategies differ across user behavior types, across both accuracy and beyond-accuracy dimensions of recommendation quality, and across the temporal-window setting that governs the disentanglement.
Physics based system forecasts LEO satellite internet quality anywhere
Reading the Sky to Forecast the Ground: Physics-Informed Link-State Forecasting for LEO Networks at Any Location
Abstract: In this paper, we introduce Gnomon, a physics-informed system that forecasts user-perceived low-Earth-orbit (LEO) downlink throughput, uplink throughput, and round-trip time (RTT) under different levels of trace availability. Gnomon's physics layer reconstructs the serving geometry and four-leg bent-pipe attenuation from public weather, orbital, routing, and licensing data. Based on what is available, Gnomon conditions on the target terminal's own history (Mode 1), measurements from nearby publicly reachable dishes (Mode 2), or the physical covariates alone (Mode 3) to predict the link state: Modes 1 and 2 share a fine-tuned time-series foundation model, while Mode 3 uses a compact boosted-tree estimator. All three modes expose a common output interface and can be selected without retraining. We evaluate Gnomon using 8,260 minutes of 1 Hz measurements collected at nine sites across five states in the U.S. We train on three sites and hold out the remaining six sites and their serving beams. On these unseen sites, the own-trace mode reduces downlink-throughput and RTT prediction error by 17% and 11% relative to the strongest published baseline and, to our knowledge, provides the first LEO uplink forecasts. The neighbor-trace mode requires no on-site hardware, while the covariate-only mode reduces downlink-throughput and RTT error by 24.6% and 78.8% relative to the only prior covariate-only forecaster. Moreover, experiments show that Gnomon provides calibrated quantile bands and improves adaptive-bitrate streaming driven over real TCP flows on replayed Starlink links.
Feature decorrelation improves balance in sequential item recommendations
A Redundancy Reduction Approach for Controllable Sequential Recommendations
Abstract: Sequential recommendation must operate under long-tailed item distributions and popularity-driven concentration, often forcing practitioners to trade short-list accuracy against long-tail exposure. In this work, we study feature decorrelation as a mechanism for shaping representation geometry in dot-product sequential recommenders, and analyze how this, in turn, affects popularity-driven concentration. We propose a decorrelation-regularized training framework that augments next-item prediction with an auxiliary redundancy-reduction term, and instantiate it with BT-SR, which uses the Barlow Twins objective. To form label-consistent positive pairs without synthetic corruptions, we pair user histories that share the same next-item target. Beyond accuracy, we provide a geometric analysis showing how decorrelation suppresses shared low-rank directions in the user representation space that can give popular items a global scoring advantage, and we introduce a bucket-based alignment concentration metric to quantify this effect. Experiments on five public benchmarks show that BT-SR consistently improves next-item ranking quality, while the decorrelation strength acts as a simple control knob that reallocates accuracy across head and tail items, enabling accuracy-exposure trade-offs. Our analysis also reveals that the impact on head-vs-tail exposure differs across datasets, reflecting interactions between decorrelation and data temporal structure.