Papers for

model developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Netflix reveals new way to measure recommendation impact accurately

Beyond Raw Engagement: A Counterfactual Observability Framework for Recommender Systems at Netflix

Abstract: Understanding the performance of large-scale recommender systems remains an underexplored challenge, especially for content creators and model developers. The raw engagement signals available to them, such as views and clicks, conflate content quality, model behavior, presentation bias, and audience reach, making it hard to attribute outcomes to the right cause. In this work, we present a general evaluation framework that enhances observability across multiple recommender systems at Netflix and demonstrate its effectiveness through several production deployments. The framework treats recommender-system observability as a counterfactual measurement problem: estimating what the recommender would have done, and what engagement would have followed, in the absence of a specific content item or model decision. We articulate three stakeholder-centered observability principles for content creators and model developers, and propose measurement methodologies covering bias reduction, relativity, and incrementality, applicable to both single-stage and cascading recommender systems and serving both audiences from a single measurement foundation.

Sat 19 SeptInformation RetrievalArtificial Intelligence
The gist
Netflix has a new way to figure out how well its recommendation system really works. Often, just looking at clicks or views doesn’t tell the full story because many things can influence those numbers. The authors created a method that estimates what would have happened if a certain show or algorithm choice hadn't been made. This helps creators and engineers understand the real reasons behind user engagement more clearly. It works across different recommendation stages and helps reduce bias in measuring results.
Open → 2609.22747v1

Backdoors in large language models separate trigger detection from control

LLM Forensics: Where Do Backdoors Hide? Localizing and Controlling Trigger Mechanisms with Sparse Autoencoders

Abstract: Even though backdoors in LLMs have been a growing concern, their inner workings are still under heavy scrutiny. Trigger-based backdoors are easy to define behaviorally, a rare input that makes the model switch to a chosen response pattern, but the mechanism between triggers and their responses is less clear. We study this mechanism in a controlled, harmless language-switching setting, where fixed trigger sequences make 1B and 8B language models continue English prompts in French or German. For this, we train sparse autoencoders (SAEs) across layers and transformer components, then compare triggered prompts with translation and pretraining controls to identify trigger-relevant feature directions. We show how SAE features separate triggered prompts from controls with near-perfect F1, but features that detect the trigger do not necessarily control the behavior. In intervention tests, attention and MLP features often fire reliably on triggered prompts, making them good detectors, but ablating them rarely suppresses the language switch and activating them rarely induces it. In contrast, residual-stream features can suppress triggered generation when ablated, and some selected features can induce target-language continuations without the trigger. In short, these token-trigger mechanisms decompose into distinct SAE feature directions, with separate features for trigger detection, residual-stream propagation, and later language tracking. This role-level decomposition is the part most likely to transfer to other trigger-based backdoors, even when the payload, layers, or circuit locations differ.

Mon 7 SeptComputation and Language
The gist
Some large language models can be tricked by special hidden phrases, called triggers, to respond in unusual ways. This paper looks inside the models to understand how these triggers cause a switch in the model’s language output, using a harmless test of switching English to French or German. The authors find that the model uses different internal features: some detect the trigger, while others control the switch in language. This separation helps explain how backdoors operate and suggests ways to find and control them in other cases.
Open → 2609.07746v1