Persona mixture models improve simulation of real human dialogue
Pretrained Persona Mixture Models and Tandem Models for Human Simulation
Computation and Language
Summary
Simulating how people talk using AI often leads to unrealistic and repetitive answers when the AI just pretends to be a certain character. The authors show that training AI on actual short conversations from real people helps the AI respond in a way that sounds more natural and diverse. They call these trained models Persona Mixture Models and find they work better than models just tuned with instructions. Combining these models with other instruction-following models makes the best predictions over different kinds of conversations.
What this means in practice
- •For chatbot developers: Create AI chatbots that simulate real human users more naturally by training on actual individual dialogue samples.$Commercial implications: Enables development of more realistic conversational agents for customer service and entertainment, improving user engagement and satisfaction.
- •For virtual assistant teams: Improve virtual assistants’ ability to adopt diverse user personas accurately for richer human-AI interaction.
Authors
Minwoo Kang, Téa Wright, Seun Eisape, Ayush Raj, Suhong Moon, Joseph Suh, Alane Suhr, David M. Chan, John Canny
Abstract
We argue here that the current dominant practice in LLM human simulation: prompting instruction-tuned assistant language models to role-play personas, is inaccurate and produces stereotyped predictions (lacking natural diversity). It has previously been shown that LLMs can be bound to personas using naturalistic, freetext dialog avoiding stereotyping. Here we show that binding can also be achieved using short, individual samples of dialog from specific people. Demographics can be added later without negative effects by simply querying the model. We use the term Persona Mixture Models (PMMs) for well-calibrated human models, currently realized as pretrained base models. We show that PMMs produce more accurate predictions than instruction-tuned models and retain more of the lexical, semantic, and pragmatic diversity found in human dialog. We measure realism and diversity of LLMs simulating human interlocutors across a diverse set of corpora spanning open-domain text, human-AI chat, and task-oriented dialogue between human speakers. However, base pretrained models can produce out-of-domain dialog and may lose some of the human's internal state over long contexts. We propose and explore tandem models which combine a pre-trained model with an instruction-tuned supervisor. Tandem models achieve the best overall accuracy and diversity in our experiments.