User simulator improves dialogue realism and persuasion accuracy

Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency

Artificial Intelligence

Summary

Simulating user conversations is important to make AI systems that chat feel real and helpful. The authors created a new simulator called TRACER that better mimics how users change their goals during a conversation and how conversations usually end. TRACER learns from real dialogues and then improves by practicing many-turn conversations to match real behaviors closely. It outperforms previous simulators, produces realistic chat flows, and helps test AI models for both conversation quality and persuasion success.

What this means in practice

  • For customer service teams: Use TRACER to simulate real customer dialogues for improving AI chatbots’ ability to handle evolving user goals and increase conversion rates.
  • For llm developers: Evaluate and optimize large language models for effective sales and support chat by testing them with the Dynamic Marketing Benchmark via simulated conversations.
  • For digital marketing companies: Deploy the simulator to create more natural and persuasive AI-driven marketing interactions that better reflect real customer behaviors.$Commercial implications: Enables AI-powered marketing software to better simulate customer conversations and improve conversion rates, supporting product offerings targeted at marketing automation.

Authors

Geng Chen, Ruotong Pan, Zhirui Yang, Qiqi He, Jiawei Chen, Zhang Yunfei, Chongyuan Chen, Minxuan Lv, Zheng Yang, Win-Bin Huang, Xiangyu Wu, Wenwu Ou

Abstract

Faithful user simulation is fundamental to building, evaluating, and improving interactive AI at scale. However, plausible individual responses do not ensure that simulated users reproduce the intent evolution and outcomes observed in real interactions. We propose TRACER, a multi-turn user simulator that explicitly models users' evolving intent and learns to align simulated behavior with real interaction trajectories. TRACER is trained in two stages: supervised fine-tuning on real user dialogues, followed by multi-turn reinforcement learning. The RL stage combines hierarchical outcome- and trajectory-level rewards with deviation-aware advantage modulation, jointly mitigating reward sparsity and credit assignment in long dialogues. On real customer-service sessions organized into reference cohorts, TRACER-7B surpasses the strongest baseline by 11.4 conversion F1, while also achieving the lowest group-level conversion-rate error and semantic trajectory distance, and generalizing to out-of-distribution scenarios. Human Turing tests yield identification accuracy close to chance, supporting the perceived naturalness of generated conversations. Building on this simulator, we further introduce the Dynamic Marketing Benchmark, which jointly evaluates persuasion effectiveness and response quality of LLMs through simulated interactions, revealing that higher response quality does not necessarily correspond to higher conversion rates.