Agent system diagnoses and improves recommendation algorithms at scale

AURA: Agentic Diagnosis and Refinement for Production Recommender Systems at Scale

Information RetrievalArtificial IntelligenceMachine Learning

Summary

Recommender systems suggest things like movies or products but sometimes fail to give good suggestions. The authors created AURA, an AI tool that looks deeply at how and where these systems make mistakes by checking large amounts of real user data. AURA can then suggest and even change the recommendation algorithms to fix problems in the code. This approach was tested on big real platforms and can be adapted for other areas like online shopping.

What this means in practice

  • For media streaming engineers: Automatically diagnose and correct failures in large-scale video and music recommendation algorithms using AI agents analyzing user engagement logs.
  • For e-commerce platform developers: Adapt AI-driven diagnosis and refinement to improve product recommendation systems in online retail environments.

Authors

SungGeun Kim, Abhinav Narain, Daniel Nemirovsky

Abstract

How and why does a recommender system fail the users it serves? Oftentimes, practitioners are left to improve their algorithms based on a combination of feedback from stakeholder teams, domain expertise, and insights from data analyses. Yet the nuances of how and where recommendations perform well or poorly for end users are difficult to discern from aggregate quantitative metrics. Whereas these metrics provide a high-level and incomplete picture, further granularity into the quality of recommendations and their patterns requires reasoning with domain understanding and objectivity, at scale. We contemplate this complex conundrum and describe a method and implementation that uses the latest AI agentic advances to provide actionable diagnoses and improvements for production recommender systems. We present AURA (Agentic Understanding and Refinement of recommender Algorithms), an end-to-end agentic system that performs qualitative evaluation at scale and can then generate improvements to our algorithms at the code level. Specialized agents read production engagement logs, from thousands of sessions to millions, and surface patterns and examples of how the recommender fails real users. The next step uses those diagnoses and context about the recommender's own code, data, and training pipeline to propose and implement refinements grounded in that codebase. We report the system design, initial tests on production data from two large consumer platforms at a major media-streaming company, safeguards, operational learnings, and early results toward a self-improving recommender system. Finally, the diagnostic gap AURA closes is not specific to streaming. The architecture is built to transfer: every domain-specific element enters through the configuration layer that already ported it between our two platforms. We map it concretely to e-commerce and online-retail recommendation.