New dialogue method improves agent responses by modeling user cognition
Proactive Dialogue Policy Optimization via Cognitive-State Transition
Artificial Intelligence
Summary
Dialogue systems often struggle to keep conversations natural and responsive over multiple turns. This paper introduces a user simulator that models how a person’s thoughts and feelings change during a conversation, helping dialogue agents give better and more consistent replies. The authors also created a new way to train dialogue policies using this simulator, which treats dialogue actions more precisely by considering both strategy and actual utterances. Their approach shows better performance on several tasks and produces more realistic user responses than some existing methods.
What this means in practice
- •For customer service teams: Improve automated chat agents to handle complex, multi-turn conversations by better modeling how users think and respond over time.
- •For game developers: Use advanced user simulators to create more realistic non-player character dialogues that adapt to player behavior and emotions across turns.
Authors
Minghui Ma, Mengqi Chen, Bin Guo, Jingqi Liu
Abstract
Proactive dialogue requires agents to continually adapt their policies to user feedback while progressing toward task objectives over multiple turns. To move beyond imitation learning on static datasets, recent approaches use user simulators to collect interactive data for policy optimization. However, many simulators do not explicitly model the evolution of user cognition, limiting the consistency and state dependence of feedback across turns. Moreover, representing each action only by a high-level strategy label overlooks the large utterance space and cannot distinguish alternative realizations of the same strategy. To this end, we jointly design a $\textbf{Cog}$nitive User $\textbf{Sim}$ulator $\textbf{(Cog-Sim)}$ and $\textbf{C}$ognitive-$\textbf{S}$tate $\textbf{T}$ransition--Driven $\textbf{P}$olicy $\textbf{O}$ptimization $\textbf{(CSTPO)}$. Cog-Sim maintains the user's cognitive and affective states and generates responses through constrained state transitions across turns, so feedback depends on both the realized utterance and the user's current state. CSTPO organizes each action as a hierarchical strategy--utterance representation: a high-level strategy label constrains utterance sampling, and utterances are optimized within each label. Sparse complete-branch sampling reuses shared dialogue prefixes and estimates separate strategy-level and utterance-level advantages, enabling fine-grained optimization at both levels. Across three tasks, Cog-Sim exhibits monotonic dose--response relationships and is preferred over prompt-based simulators for naturalness. CSTPO improves Qwen3-14B's performance to a level comparable to that of GPT-5.5-based planning methods.