Llms reveal and adjust internal user beliefs to guide responses

User Model Extraction via Belief Self-Distillation

Machine LearningComputation and Language

Summary

Large language models (LLMs) guess information about the person they are talking to and change how they reply based on these guesses. The researchers created a new method called Belief Self-Distillation (BSD) that lets us see and change what the model believes about its user from regular conversations, without needing extra labels. This method helps test how changing these beliefs affects what the model says, for example, why it refuses some requests. They also found that different independently trained models organize their user beliefs in very similar ways. This work helps make the internal thinking of LLMs more understandable and controllable.

What this means in practice

  • For ai developers: Build safety features that adjust model responses by directly changing inferred user beliefs rather than just input prompts.
  • For chatbot designers: Create chatbots that transparently reveal and control their assumptions about users to improve interaction quality and trust.

Authors

Ali Holmov, Yiran Huang, Kirill Bykov, Zeynep Akata

Abstract

Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stronger interventions than matched hidden-state steering. Crucially, we find that refusal depends not only on the request, but on the model's inferred user intent: changing this belief alters refusal while holding the request fixed. We further uncover a striking cross-model regularity: independently trained LLMs converge on a shared geometry for representing their users. Together, these results reveal implicit user models as readable and causally writable internal states with direct implications for AI safety, shaping how models condition safety decisions on whom they believe they are interacting with.