Papers for

virtual assistant teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Language agents that communicate effectively with fewer words

Towards Communication-Efficient Social Intelligence in Language Agents

Abstract: Socially intelligent language agents must negotiate, coordinate, and resolve conflicting preferences while respecting the time and attention of both participants. Balancing these demands is challenging because agents must convey enough to address a partner's constraints and advance their goals without adding words that do not help the interaction. In this paper, we propose Teacher-Assisted Communication Training (TACT) to improve social goal attainment while reducing communication cost, making interactions with agents more productive and less demanding. We first characterize communication efficiency in terms of action strategy and expression, whose effects extend beyond the current utterance to the partner's response and subsequent exchanges. We design TACT to revise student-generated actions, test the revisions through partner responses, and distill useful feedback into the student. An expression specialist removes unnecessary detail while preserving the intended action, while a strategy specialist proposes alternatives that may better address the partner's constraints. To determine which revision helps, TACT samples a partner response for each candidate and selects a teacher reference by balancing local goal support against action-token cost. That reference guides on-policy distillation on the student's own generation prefixes, allowing the student to act independently at deployment. We evaluate TACT on SOTOPIA and AgentSense. On SOTOPIA, it achieves the highest Goal among the evaluated methods on All and Hard while using substantially fewer target tokens than SFT+SDPO. On AgentSense, it improves goal success over the initial student while reducing target tokens and interaction messages.

Mon 28 SeptComputation and Language
The gist
Socially smart language agents need to cooperate by sharing just the right amount of information, so conversations stay useful without being long or confusing. The authors created a method called Teacher-Assisted Communication Training (TACT) that teaches these agents to pick their words and actions carefully, helping them meet goals while using fewer words. TACT works by checking different ways of saying something, guessing how the partner might respond, and then learning which version works best with the smallest cost in words. The authors tested TACT in two settings and found it improved success rates while cutting down on unnecessary communication.
Open → 2609.35749v1

Models often fail to forget wrong user claims after reset

Reset Is Not Recovery: Evaluating Recoverability from False Conversational Context via Sycophancy Hysteresis

Abstract: Grounded language models are usually evaluated by adding relevant context, but multiturn dialogue also contains unsupported user claims that may contaminate later factual answers. We study post-pressure recoverability: whether a model returns to clean-context behavior after a user repeatedly advocates a wrong answer and then withdraws that pressure. We introduce a recovery-after-pressure protocol for multiple-choice factual dialogue and measure sycophancy hysteresis, the residual probability assigned to the user-advocated wrong answer relative to a clean-context counterfactual. Across seven instruction-tuned open-weight models and two factual benchmarks, ordinary reset often reduces but does not erase pressure-induced bias. History preserving repairs such as user retraction, system reset, and self-verification recover only 2-3/14 model-dataset pairs under the strict clean-restoration diagnostic, whereas operations that change the effective context are substantially more reliable; the two conditions that remove the pressure-bearing history entirely, fresh-context deletion and context truncation, recover 14/14. In an oracle trusted-evidence condition across fourteen model-dataset pairs, preserving the pressure-bearing history while adding benchmark-derived trusted evidence increases accuracy from 0.368 to 0.929, while wrong-answer following falls from 41.2% to 4.3%. Controls show that the effect is not explained by dialogue length, repeated confidence, plausible distractors, mere false-answer mention, or option-label inertia. These results suggest that faithful grounded dialogue requires evaluating which prior context should be treated as evidence and which should be removed or quarantined before answering.

Sun 27 SeptComputation and Language
The gist
Sometimes AI chat models are told wrong facts by users during a conversation, and later they try to correct themselves. This paper studies whether these models can truly recover and stop believing the wrong information, once the user admits the mistake. The authors found that simply resetting the conversation often doesn't fully remove the wrong influence. Only completely removing or truncating the chat history reliably restores the model to giving correct answers. Adding trusted evidence to the conversation greatly improves the model’s accuracy even if previous wrong claims remain.
Open → 2609.33672v1

Persona mixture models improve simulation of real human dialogue

Pretrained Persona Mixture Models and Tandem Models for Human Simulation

Abstract: We argue here that the current dominant practice in LLM human simulation: prompting instruction-tuned assistant language models to role-play personas, is inaccurate and produces stereotyped predictions (lacking natural diversity). It has previously been shown that LLMs can be bound to personas using naturalistic, freetext dialog avoiding stereotyping. Here we show that binding can also be achieved using short, individual samples of dialog from specific people. Demographics can be added later without negative effects by simply querying the model. We use the term Persona Mixture Models (PMMs) for well-calibrated human models, currently realized as pretrained base models. We show that PMMs produce more accurate predictions than instruction-tuned models and retain more of the lexical, semantic, and pragmatic diversity found in human dialog. We measure realism and diversity of LLMs simulating human interlocutors across a diverse set of corpora spanning open-domain text, human-AI chat, and task-oriented dialogue between human speakers. However, base pretrained models can produce out-of-domain dialog and may lose some of the human's internal state over long contexts. We propose and explore tandem models which combine a pre-trained model with an instruction-tuned supervisor. Tandem models achieve the best overall accuracy and diversity in our experiments.

Fri 18 SeptComputation and Language
The gist
Simulating how people talk using AI often leads to unrealistic and repetitive answers when the AI just pretends to be a certain character. The authors show that training AI on actual short conversations from real people helps the AI respond in a way that sounds more natural and diverse. They call these trained models Persona Mixture Models and find they work better than models just tuned with instructions. Combining these models with other instruction-following models makes the best predictions over different kinds of conversations.
Open → 2609.22607v1

Automated views improve memory retrieval in long conversations

AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory

Abstract: Long-term memory is essential for large language model (LLM) agents to maintain consistency and personalization over extended interactions. Existing memory systems typically rely on fixed granularities or static schemas, but these designs struggle when heterogeneous information, such as preferences, events, constraints, and temporal updates, is embedded in a single mixed representation. The resulting semantic interference makes top-K retrieval sensitive to noise and often leaves relevant evidence poorly ranked. We present AutoViewMem, a data-driven framework that organizes long-term conversational memory into self-configuring, low-overlap semantic views before indexing. AutoViewMem discovers candidate views from interaction traces, selects a compact complementary view set, and uses these views to guide write-time structured extraction of provenance-grounded memories. This representation-first design moves semantic disentanglement from retrieval time to write time, allowing standard top-K similarity search to retrieve focused evidence without explicit routing or iterative retrieval. We further apply offline consolidation to improve memory compactness and consistency. Experiments on the LoCoMo and PersonaMem benchmarks, under both Qwen3-8B and Qwen3-14B backbones, show that AutoViewMem improves long-horizon question answering and personalization over strong memory baselines while preserving a simple inference pipeline.

Fri 18 SeptArtificial Intelligence
The gist
Remembering details accurately is hard for AI chatbots during long talks. The authors show a way to organize the chatbot’s memories into separate groups based on topics, so it doesn’t mix up different types of information. This helps the AI find the right facts more easily when answering questions or personalizing chats. The system they created works better than older methods and keeps things simple when the AI responds.
Open → 2609.21940v1

GLARE improves forecasting of conversation flow in real meetings

GLARE: Generative Learning via Adversarial Reward Estimation For Social Dynamics Forecasting

Abstract: Meeting continuation requires tracking the agenda, speaker roles, participant intentions, and disagreement across long multi-party discussions. We introduce the Meeting Dynamic Forecasting Benchmark (MDFB), constructed from 2,207 real-world meetings and 24,794 future-facing queries. Given a transcript prefix and an active question, a model generates a plausible multi-turn continuation in one call. We evaluate utility---progress toward the question---and human-likeness---plausible conversational flow and role consistency---without requiring exact reproduction of the observed future. We further present GLARE, an adaptation of adversarial imitation learning to conditional language generation. A discriminator ranks the observed continuation above samples from the current actor, and its score supplies a KL-regularized policy reward; retraining on current-policy negatives allows the reward landscape to evolve with the actor. GLARE attains average human-evaluated win rates of 0.66 on utility and 0.70 on human-likeness, outperforming SFT and SPIN while remaining below the observed human continuation. We also demonstrate MDFB as a social reasoning arena for comparing general-purpose models, including closed-source systems, through reference-assisted judgments. Together, these studies illustrate the benchmark's use for both task-specific learning and output-based evaluation of meeting behavior.

Thu 10 SeptArtificial Intelligence
The gist
Following the flow of real meetings is hard because people talk, have roles, and change their minds over time. To help with this, the authors created a big set of meeting recordings and questions about what happens next, called MDFB. They also built a method called GLARE that teaches computers to guess what might happen later by comparing guesses to real meeting parts and learning from that. GLARE does better than earlier methods at making useful and natural-sounding meeting continuations.
Open → 2609.12165v1