Papers for

virtual assistant designers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Reactive audio-driven model improves listener facial motion in robots

REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening

Abstract: Generating responsive listener facial motion is an important task for embodied conversational AI. Two modeling challenges are central: accounting for the timing of speaker cues while maintaining continuity with the listener's ongoing motion, and capturing locally variable facial events alongside the overall motion trajectory. Listener responses may follow preceding cues with a temporal lag, while brief expressions and blinks introduce variation that is difficult to predict deterministically. These challenges motivate a framework that combines history-aware temporal alignment with stochastic expression refinement. We propose REALM (Reactive Embodied Audio-driven Listening Model), a coarse-to-fine framework for audio-driven reactive listening. A Reactive Gated Speaker-Listener Fusion module combines listener motion history with speaker audio through a delay-centered attention prior and adaptive gating. A coarse decoder predicts a base motion trajectory, which is augmented by audio-conditioned stochastic residuals in the expression subspace while retaining the coarse pose parameters. Evaluations on ViCo and L2L show improvements over the evaluated baselines across multiple motion-quality metrics. Additional analyses examine delay sensitivity, gate behavior, and blink dynamics. Finally, deployment on an Ameca humanoid robot and a perceptual user study demonstrate the applicability of the generated behavior to physical embodiment. Code: https://github.com/lipzh5/REALM Demo: https://youtu.be/Tf5mpd5S8VQ

Sun 27 SeptRobotics
The gist
Generating natural facial reactions for virtual or robotic listeners during conversations is tricky because listener responses depend on timing cues and subtle expressions like blinks. The authors created a system called REALM that first predicts broad listener facial movements and then adds small, random expression details triggered by the speaker's audio. Their model helps robots and avatars react more naturally to the speaker, making conversations feel more lifelike. They tested it on datasets and a humanoid robot to show it works better than previous methods.
Open → 2609.33095v1

LLM agents struggle with changing user intentions in conversations

When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents

Abstract: LLM agents often operate over multi-turn interactions in which user intent changes before execution. We study intent drift: the failure mode in which superseded parts of the user's intent continue to influence the final answer or tool action. We introduce IntentFlux, an executable benchmark that converts verifiable tasks into dialogues with controlled intent changes while preserving their original graders. In a 627-case calibration, mean task score falls from 0.476 to 0.384 as dialogues contain more superseded and withdrawn information. Across eight models, the rate of fully correct solutions is significantly lower when the same final task must be recovered from an evolving dialogue rather than given directly in a single turn. We further introduce StateForge, which explicitly maintains the active requirements before generation. On General-Test, it improves mean task score from 0.367 to 0.467. Providing the ground-truth final state improves performance further but still does not recover single-turn performance, indicating that state-estimation errors explain only part of the gap. These results establish intent drift as a measurable multi-turn failure mode and explicit state maintenance as a partial mitigation.

Sat 26 SeptComputation and LanguageArtificial Intelligence
The gist
Sometimes when people talk with AI tools, they change their minds partway through. This causes the AI to get confused and make mistakes based on old information. The authors created a way to measure how much this happens and built a tool to keep track of what the user really wants at each step. Their tool improved AI accuracy but didn’t completely fix the problem, showing it’s a tricky challenge.
Open → 2609.32520v1

Speech drives realistic 3D facial animation with visible lip movements

Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation

Abstract: Recent progress in speech-driven 3D facial animation has improved vertex-level reconstruction quality, but speech-consistent visible articulation remains difficult. This is because speech production follows structured and constrained articulators' coordination and the mapping from acoustics to motion is inherently one-to-many. Motivated by the structured patterns of visible articulation, we propose a novel articulation-aware framework that models visible speech through directional articulatory motions and composes them into surface-consistent 3D facial motion. To represent visible articulation with three directional articulatory motions, spreading, opening, and protrusion, we propose a Speech--Articulatory Memory (SAM) that captures the correspondence between speech and these motions under phonetic context through retrieval and decoding based on a key-value memory structure. Then, a Topology-aware Articulatory Composition (TAC) integrates the predicted directional articulatory motions under mesh topology to produce surface-consistent 3D facial motion. Experiments on VOCASET and TFHP show that our method achieves state-of-the-art performance on standard reconstruction metrics and improves visible articulatory distance and velocity errors for lip articulation, while a user study confirms clear preference in lip sync and realism.

Thu 24 SeptGraphicsMachine Learning
The gist
It is hard to create 3D facial animations that match spoken words perfectly because many mouth shapes can produce the same sounds. The authors built a new system that understands mouth movements in three directions—spreading, opening, and sticking out—to better match speech. Their method uses a special memory to link sounds to these movements and combines them in a way that keeps the 3D face model smooth. Tests showed this makes lip animations more accurate and natural looking than previous methods.
Open → 2609.30517v1

Turn-taking in dialogue models improves by using speaker intent timing

Neither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models

Abstract: Benchmarks for full-duplex spoken dialogue models score turn-taking with binary fixed-window rules that reward immediate response or silence by completeness of the prior turn. We argue that the appropriateness of a response offset, whether delayed silence or anticipatory overlap, is conditional on the speaker's latent intent, identifiable only from that speaker's behavior. We introduce TACT, a benchmark of 9,728 episodes and 73.2 hours from five dyadic corpora; each episode carries dialogue history, a per-speaker memory profile, and an annotator-derived posterior over six intent classes. Scoring replaces binary windows with a strictly proper threshold-weighted continuous ranked probability score whose weights are intent-conditioned timing kernels fitted to human floor-transfer-offset distributions, proving boundedness, consistency, and binary reduction. Across eleven systems the best model reaches 0.47 against a human topline of 0.86, is nearly invariant to speaker profiles, and TACT agrees with human judgments at Spearman 0.81 versus 0.46 for binary metrics.

Wed 23 SeptComputation and LanguageArtificial IntelligenceHuman-Computer Interaction
The gist
Current methods to judge when one person should speak after another in conversations are too simple, only allowing quick replies or silence. The authors show this isn’t enough because people sometimes wait or talk over each other depending on their intentions. They created a new benchmark called TACT with real dialogue data and detailed information about what speakers intend to do. This benchmark scores turn-taking by comparing timing patterns to those of humans, improving agreement with human judgments. Their tests show models can get better at understanding conversation flow beyond simple rules.
Open → 2609.27372v1

Gesturefar enables real-time natural gesture generation from streaming speech

GestureFAR: Streaming Co-Speech Gesture Generation with Flow Autoregression

Abstract: Generating natural co-speech gestures from streaming speech is essential for embodied conversational agents, where motion must be produced while a user is still speaking. Recent streaming gesture systems make online generation possible by autoregressing over discrete motion tokens, but this design compresses high-dimensional continuous motion into finite codebooks and can limit the realism and diversity of generated gestures. To preserve both causality and continuous expressiveness, we propose \textbf{GestureFAR}, a flow-autoregressive framework for streaming co-speech gesture generation. First, GestureFAR autoregresses over causal continuous motion latents, using a transformer to model streaming audio-motion context and a per-token flow-matching head to sample the next latent from a continuous distribution. Second, we introduce a head-only flow distillation strategy that freezes the causal backbone and distills the multi-step per-token flow head into a single network evaluation using consistency and distribution-matching objectives. This keeps the model token-causal while removing the main latency bottleneck for live interaction. Experiments on BEAT2 show that GestureFAR significantly improves the quality--latency trade-off among streaming-capable methods, preserving strong gesture quality while enabling real-time token-causal generation. Project Page: https://andypinxinliu.github.io/GestureFAR

Fri 18 SeptComputer Vision and Pattern RecognitionGraphicsHuman-Computer Interaction
The gist
Making digital characters move their hands naturally while talking is tricky, especially when the speech is still happening. Previous methods simplified complex hand motions into limited building blocks, which made the gestures less lifelike and varied. The authors created GestureFAR, a new way that predicts smooth, continuous hand movements directly from what is being said, without waiting for the speaker to finish. This approach lets virtual agents produce more natural gestures instantly as speech flows.
Open → 2609.21576v1