Semantic user profiling methods compared for streaming recommendations

When LLM-Based User Profiling Adds Value in Production Streaming Recommendation

Information Retrieval

Summary

Online services often suggest content based on what users liked before, but creating user profiles from past actions varies in complexity and cost. The authors compare two main ways to build these profiles: one that averages numerical representations of items, and another that uses large language models (LLMs) to write natural descriptions of user interests. They tested these methods on a real streaming service dataset, looking at how well each works depending on the user’s behavior and the timing of actions. Their findings help decide when the expensive LLM-based profiling is worth using compared to simpler methods.

What this means in practice

  • For streaming service engineers: Choose user profiling methods that balance cost and accuracy based on user behavior and content timing to improve recommendation systems.
  • For e-commerce platform developers: Implement semantic user profile strategies to tailor product recommendations depending on recent versus historical customer interactions.

Tested on one dataset.

Authors

Milad Sabouri, Neeraj Sharma, Sardar Hamidian, Shaghayegh Agah

Abstract

Personalized recommendation depends critically on how user representations are constructed from historical behavior. Two paradigms have emerged for constructing semantic user profiles in content-based recommendation. First, aggregate methods derive user representations as numerical aggregates of semantic item embeddings. Second, LLM-based methods generate natural-language summaries of user preferences and encode them through a text encoder. Each paradigm can be combined with temporal disentanglement of recent versus historical behavior. LLM-based profile generation is significantly more expensive than aggregate approaches, raising the question of when this additional cost is justified. We present a systematic comparison of four semantic user-profiling strategies, factorially crossed across representation type and temporal handling, evaluated on a real-world production dataset. The comparison reveals how these strategies differ across user behavior types, across both accuracy and beyond-accuracy dimensions of recommendation quality, and across the temporal-window setting that governs the disentanglement.