Rubric aware learning improves personalized answers from large language models

Rubric-Aware On-Policy Self-Distillation for LLM Personalization

Computation and Language

Summary

Getting large language models (LLMs) to give answers that match exactly what a specific user wants is tricky. The authors created a new way called GRASP that uses detailed user instructions (rubrics) to teach the model how to generate better personalized responses, down to each word it picks. They also check the quality of these instructions to avoid bad teaching moments. Their method showed better results than previous ones on a personalized question answering test.

What this means in practice

  • For chatbot developers: Use rubric-aware token-level learning to improve chatbots that better follow specific user preferences in their answers.
  • For customer support teams: Enhance automated support tools to provide personalized answers that closely match detailed customer needs.

Authors

Yilun Qiu, Xiaoyan Zhao, Chengbing Wang, Cilin Yan, Rui Zu, Wanyang Zhang, Xiaolong Jiang, Jiayin Cai, Yang Zhang

Abstract

LLM personalization aims to generate responses aligned with individual users' preferences and needs. User-specific rubrics make these expectations explicit, providing direct supervision on what a satisfactory answer should cover. Existing rubric-guided approaches, however, exploit such guidance only at a coarse granularity, either by using rubrics to supervise the prediction of relevant aspects for subsequent generation or by reducing aspect coverage to a single response-level reward for reinforcement learning. This leaves a gap between specifying what a personalized answer should contain and teaching the model how to generate it. To bridge this gap, we propose GRASP, a rubric-aware on-policy self-distillation framework for LLM personalization that turns user-specific rubric aspects into fine-grained, token-level supervision. Specifically, GRASP pairs a rubric-free student with a rubric-informed teacher that additionally receives the target user-specific rubrics. By aligning their next-token distributions along on-policy trajectories generated by the student, GRASP transfers the teacher's rubric-conditioned guidance into the student, translating user-specific semantic requirements into dense token-level supervision. Since rubric-informed teachers can still produce inadequate supervision, we further introduce Rubric-based Teacher Validation (RTV), which retains only instances where the teacher sufficiently covers the target aspects, improving both supervision quality and training efficiency. Experiments on the LaMP-QA benchmark for personalized question answering demonstrate that GRASP achieves state-of-the-art performance across multiple backbones, supporting the effectiveness of rubric-guided token-level supervision for personalization. To ensure reproducibility, our code is available at https://github.com/SnowCharmQ/GRASP.