On policy self distillation improves persona consistency in dialogues
OSPD: On-Policy Self-Distillation for Persona-Consistent Dialogue
Artificial Intelligence
Summary
Keeping a character's personality consistent over a conversation with AI is hard. The authors created a new method called OSPD that lets the AI learn from itself, using different details about the character for teaching and speaking. This approach helps the AI stay more true to the character's personality without needing extra outside models or rewards. Tests show it beats other methods at making dialogues that sound like a consistent character.
What this means in practice
- •For chatbot developers: Create chatbots that maintain personality traits more consistently over multiple turns using on-policy self-teaching techniques.$Commercial implications: Enables building advanced conversational agents with stable character personas, improving user experience for virtual assistants and entertainment bots.
- •For game narrative designers: Develop NPCs with more consistent role-playing dialogue that better follows character profiles without extra external models.
Authors
Rui Xu, Yikai Zhang, Aili Chen, Zicheng Zhao, Xu Yinghui, Libo Wu
Abstract
Maintaining persona consistency across multi-turn dialogues remains a core challenge for role-playing language models. Off-policy distillation from external teachers incurs distribution mismatch that compounds across dialogue turns, while reinforcement learning struggles with reward ambiguity inherent in subjective persona fidelity. We propose OSPD, an on-policy self-distillation framework where the same model serves as both teacher and student under asymmetric information: the teacher receives a complete character profile while the student sees only a brief summary, and the student generates trajectories from its own policy. We find that teacher confidence in role-playing dialogue exhibits a bimodal structure---sharply peaked at character-critical tokens yet diffuse at generic utterances---and introduce role-aware divergence switching to match this structure. A progressive trait masking curriculum further forces staged internalization of character knowledge along semantic dimensions. Experiments on CharacterBench, CharacterEval, and SocialBench show that OSPD substantially improves persona consistency over supervised fine-tuning and multi-turn RL baselines, without requiring any external teacher or reward model.