Papers for

voice assistant developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Full duplex voice agents adapt speech while others talk

Continue, Adapt, or Yield: In-Turn Adaptation to Overlapping Speech in Full-Duplex Agents

Abstract: Full-duplex evaluation often emphasizes whether an agent keeps speaking or stops. That binary cannot express a third response humans use routinely: continuing to speak while incorporating what the listener just contributed. The contribution may be a missing word, a correction or a clarification. We introduce Duplex Cue, an evaluation of this \emph{in-turn adaptation} in full-duplex voice agents. Duplex Cue separates listener intent (backchannel, collaboration, or interruption) from speaker behavior: continuing unchanged, adapting within the turn, or yielding. Adaptation includes acknowledgment as well as content revision. In a single-model case study using 300 human-confirmed cues from unscripted English conversations, we compare recorded human responses with PersonaPlex continuations generated while replaying the listener's audio. We retain 208 pairs with the ongoing speaker active at cue onset and a scorable response in each condition. On the 66 collaborative pairs, recorded speakers adapt in 68.2\% of cases, compared with 34.8\% for PersonaPlex. The model otherwise continues unchanged (42.4\%) or yields (22.7\%). These findings show why evaluating natural voice interaction requires measuring how an agent responds to a listener's contribution as well as whether it keeps speaking.

Fri 11 SeptComputation and LanguageSound
The gist
Talking with a voice-controlled assistant can be tricky when people speak at the same time. The authors point out that humans don’t just stop or continue speaking when interrupted; they often adjust their words on the fly to include what the other person said. They developed a new way, called Duplex Cue, to measure how well AI voice agents do this kind of in-turn adaptation. Testing with a model named PersonaPlex showed it adapts less than humans, suggesting current AI still struggles with natural back-and-forth talking.
Open 2609.13117v1

Voice agents struggle to join group conversations naturally

MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant

Abstract: Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-end architectures. However, while recent benchmarks extensively evaluate dyadic interactions and passive audio comprehension, they largely overlook a prevalent real-world scenario: multi-party conversations. Evaluating agents in these settings is fundamentally more challenging than in dyadic interactions due to the exponentially greater conversational complexity. For voice agents to integrate seamlessly into human group dynamics, they must not only generate contextually appropriate responses but also demonstrate a nuanced understanding of open turn-taking. To address this gap, we introduce Multiparty Bench (MP-Bench), the first benchmark specifically designed to objectively evaluate conversational speech systems as active participants within multi-party contexts. MP-Bench assesses agent behavior along two primary dimensions: turn-taking awareness and response appropriateness. Additionally, we incorporate comprehension-based question-answering tasks as a complementary evaluation. By benchmarking 12 voice agents, we find that real-time voice agents stay at or below 22% on multiparty comprehension and remain near chance on multiparty turn-taking, exposing an open challenge for real-time voice agents under multiparty scenario.

Fri 11 SeptArtificial IntelligenceComputation and Language
The gist
Voice agents today are good at talking with one person, but they find it very hard to join conversations with multiple people. The authors created a new test called MP-Bench to see how well these agents understand when to speak and how to respond correctly in group talks. They tested 12 popular voice agents and found that they perform poorly, almost guessing when to speak and understanding very little of the conversation in real time. This shows that more work is needed for voice assistants to work well in group settings.
Open 2609.13076v1

Steerable full duplex speech models improve conversational control and timing

SteerDuplex: Steerable Duplex Speech Dialogue Models

Abstract: Full-duplex spoken dialogue models support low-latency turn taking, interruption handling, and backchanneling, yet a key capability remains underexplored: steerability, the ability to reliably shift conversational behavior along attributes such as tone, persona, speaking rate, and voice style in response to user instructions. We introduce a taxonomy of text- and audio-based steerability that identifies substantial gaps in current full-duplex models. To address this gap, we introduce SteerDuplex, a Moshi-based full-duplex speech model fine-tuned on natural conversations and synthetic dialogues targeting instruction following, vocal delivery, reasoning, and duplex interaction. We further apply two-stage reinforcement learning (RL) with hybrid rewards, combining verifiable interaction checks and judge-based semantic feedback to improve timing and response continuity. To evaluate full-duplex spoken steerability, we introduce SteerBench, a benchmark with 390 spoken prompts and 1,067 human-authored binary audio and text rubrics spanning tone, persona, style/accent, and speed/length. On SteerBench, supervised training improves audio-steering average pass rate by 44.5 percentage points over the strongest evaluated open baseline. On Audio MultiChallenge, task average pass rate improves by 7 points over its strongest evaluated open baseline. RL further raises source-clean interruption response from 72.5% to 82.5% and reduces synthetic pause barge-in from 26.5% to 9%. Steering and aggregate task scores remain comparable or higher, while reward probes reveal reward hacking through incomplete responses. Our model and benchmark support systematic research on spoken steerability, with reward analysis showing why timing gains must be evaluated alongside response completeness.

Fri 11 SeptArtificial IntelligenceComputation and Language
The gist
People want voice assistants and dialogue systems to talk more naturally and follow instructions about how to speak, like changing tone or speed. The authors found that current full-duplex speech models, which can listen and talk at the same time, were missing this ability to change their speaking style reliably. They created SteerDuplex, a model trained to adjust conversation style and timing based on user instructions, and tested it on a new benchmark called SteerBench. Their model showed big improvements in controlling voice style and handling turn-taking smoothly, though some issues remain with incomplete responses.
Open 2609.12623v1

TokenMapper enables direct speech token translation between models

TokenMapper: A Step Toward Interoperable Speech Token Translation

Abstract: Neural audio codecs discretize speech into token sequences, but the resulting token spaces differ in vocabulary and codebook structure, preventing direct communication across models. This limitation affects applications such as conversational voice agents and speech to speech translation systems where multiple speech models must interact. As a result, transferring information between speech systems typically requires decoding to waveform audio and re-encoding with a second tokenizer, increasing latency and introducing potential information loss. To address these limitations, we present TokenMapper, a direction aware framework for direct token to token translation between heterogeneous speech tokenizers in the discrete domain. TokenMapper supports structurally mismatched token spaces, including mappings between single codebook and multi codebook representations, under a shared effective token rate. Experiments on GLM-4-Voice, MiMi and DualCodec show consistent cross model performance. Specifically, translation WER approaches native reconstructions within 2.5-6.8% absolute WER, human MOS for TokenMapper outputs ranges from 2.29 to 4.39, following the same direction level trends as UTMOS and end to end latency is reduced by 4.8-94.5% relative to waveform bridging, reaching up to 972 ms per utterance. These results provide a practical step toward cross model speech token interoperability without intermediate waveform reconstruction.

Fri 11 SeptMachine LearningArtificial IntelligenceSound
The gist
Different speech processing systems break down voice into sequences of tokens, but these tokens don’t match up between models, so communication between systems is slow and can lose information. The authors created TokenMapper, a tool that translates these tokens directly between speech models without turning them back into sound first. This method reduces delays and improves accuracy compared to traditional ways. Their tests show the translations are close to the original quality while being faster.
Open 2609.12563v1

Adaptive layer reduces false wake ups in voice assistants

Not All Speech Is Intent: Adaptive Self-Correcting Inference Layer for Post-ASR False Wake-Up

Abstract: False wake-up activations remain a persistent challenge in conversational AI. Speech phonetically similar to a device's wake word can produce a syntactically valid and semantically coherent ASR transcript that the assistant incorrectly executes. Most existing systems make a single intent decision in isolation, without a mechanism to learn from recurring errors over time or adapt to individual users through personalized learning. We introduce the Feedback-Driven Adaptive Self-Correcting Inference Layer (ASCIL), a complementary post-ASR correction framework that re-evaluates wake-up intent before response generation by fusing acoustic embeddings, linguistic cues, device context, and patterns from past misclassifications. ASCIL interprets implicit signals, including hesitation, disengagement, and silence, and explicit signals, including cancellation and repetition, as automatically inferred, noisy behavioral indicators of potential misclassification. These signals drive online pattern updates without manual annotation, whereas the intentional/unintentional reference labels used for offline evaluation are human-annotated. It generalizes from prior errors, applies corrective adjustments at inference time, and continuously updates in parallel with natural-language execution. Evaluated on a proprietary dataset of 3,667 interactions with human-annotated intentional/unintentional reference labels spanning 14 acoustic and contextual conditions, ASCIL achieves 54.27% relative error reduction on a session-disjoint subset constructed from baseline failures, and up to 24.39% relative error reduction at threshold 0.90 on the issue-tagged evaluation slice. These gains are achieved while improving intentional acceptance rates, with a median added latency below 60 ms in the reported benchmark.

Fri 11 SeptComputation and LanguageArtificial Intelligence
The gist
Sometimes voice assistants mistakenly think they’ve been called when a person actually didn’t mean to speak to them, causing annoying errors. The authors developed a new system called ASCIL that listens again after the assistant is triggered and uses clues like hesitation or silence to check if the wake-up was intentional. It learns from past mistakes and adjusts itself over time without needing people to label the errors manually. Their tests showed ASCIL decreased these false alarms by over half while keeping the assistant’s responses quick and accurate.
Open 2609.12469v1

Speech enhancement improves with one-step dual latent drifting approach

DriftSE: Speech Enhancement with Generative Drifting

Abstract: We propose DriftSE, a novel one-step generative framework for speech enhancement formulated as a latent distribution equilibrium problem. During training, the drifting field aligns the generator's pushforward distribution with the clean speech manifold through drifting in a latent domain. During inference, the drifting process is discarded, enabling one-step generation. We establish that its enhancement quality depends fundamentally on the choice of latent representation. Semantic latents preserve phonetic structure but fail to capture physical acoustic cues, whereas acoustic latents reconstruct the physical signal but risk linguistic hallucination. Therefore, we introduce dual-latent drifting, performing parallel drifting in both semantic and acoustic latents to simultaneously preserve phonetic intelligibility and acoustic fidelity. Additionally, we demonstrate that DriftSE enables fully unpaired training by aligning latent distributions rather than exact point-wise targets. Consequently, DriftSE facilitates cross-dataset learning in the absence of paired noisy-clean samples. Moreover, DriftSE exhibits broad architectural flexibility across different generator backbones. Extensive evaluations on additive denoising and convolutive dereverberation demonstrate robust one-step enhancement across both offline and real-time causal settings. Notably, DriftSE achieves state-of-the-art word error rates across all four evaluated datasets while strictly operating at 1 NFE. Code and audio examples are available online.

Thu 10 SeptSoundArtificial Intelligence
The gist
Cleaning up noisy speech recordings is important for clearer communication and better voice recognition. The authors created a method called DriftSE that improves speech clarity in just one step by using two kinds of hidden information about sounds: one that focuses on meaning and one that focuses on the actual physical sound details. This approach works well even without examples of noisy and clean speech pairs during training and can be used in real-time applications. Tests show that DriftSE leads to better word recognition accuracy than previous methods.
Open 2609.12252v1

Neural models improve distant speaker diarization in noisy conditions

Neural Multichannel Distant Speaker Diarization with Heavy-tailed Source Separation Model

Abstract: Distant speaker diarization remains challenging due to difficult acoustic environments, varying numbers of speakers and overlapping speech. Model-driven methods are proposed to exploit the speech source features in multi-channel recordings that help diarization. This paper generalizes a neural model that jointly learns to perform blind source separation and diarization over speech mixtures (neural FCASA) with heavy-tailed models. The popular Gaussian distribution has been applied for variance modeling in the original source separation model, which we replace with two families of heavy-tailed models (Leptokurtic Generalized Gaussian distribution and Student's t distribution) to better capture the heavy-tailedness in speech signals. Thanks to the Gaussian scale mixture model, we are able to unify the proposed method and the original one under the same form of learning objective. Our experiments show consistent large improvements in Diarization Error Rate (DER) and Jaccard Error Rate (JER) compared to the baseline on various corpora.

Thu 10 SeptSoundArtificial Intelligence
The gist
Separating and identifying who is speaking in a room with multiple people talking at once and from a distance is hard. The authors improved a method that uses multiple microphones and advanced math models to better separate voices and label who spoke when. They replaced the usual assumptions about voice signal patterns with more flexible ones that better fit real speech. Their experiments show this change reduces mistakes in identifying speakers.
Open 2609.12154v1

Speech language models improve reasoning accuracy with real time self correction

RetroThinker: Enabling Retrospective Thinking in Speech LLMs

Abstract: Speech large language models (SpeechLLMs) offer reduced latency and retain paralinguistic nuances that are typically lost in cascaded automatic speech recognition (ASR) and text-based LM architectures. However, they continue to lag behind text-only LLMs on complex reasoning tasks, while real-time spoken interaction imposes strict latency constraints. Although prior works employ Chain-of-Thought (CoT) and concurrent reasoning to enhance reasoning capabilities without inducing prohibitive delays, an inherent accuracy-latency trade-off persists. In this paper, we investigate whether a streaming SpeechLLM can dynamically revise its reasoning traces on the fly. We introduce RetroThinker, a multi-stage post-training framework that equips the Moshi model to self-verify and forward-correct CoT steps during inference. RetroThinker combines supervised fine-tuning (SFT) on curated retrospective thinking data with length-based direct preference optimization (DPO) to optimize retrospective during early reasoning (i.e., reasoning concurrently while the user speaks). Evaluated on the GSM8K benchmark, RetroThinker significantly improves the accuracy-latency trade-off over non-retrospective baselines, achieving an 11% absolute accuracy gain at a comparable latency.

Thu 10 SeptArtificial IntelligenceComputation and Language
The gist
Speech language models can understand spoken words faster and keep voice details better than converting speech to text first. But they are not as good as text-based models at solving tricky problems quickly. The authors created RetroThinker, which lets a speech model check and fix its own thinking steps while listening. This approach helps the model be more accurate without slowing it down much. Tests showed RetroThinker improved problem-solving scores by 11% with similar speed.
Open 2609.11864v1

Differential privacy method improves federated speech model training accuracy

Component-Aware Differential Privacy for Federated Multilingual Speech-LLMs

Abstract: Per-layer differential privacy (DP) clipping improves gradient fidelity in federated learning by allocating per-matrix clipping budgets proportional to parameter count. We show that this recipe breaks for speech large language models (speech-LLMs), when the acoustic encoder and the language decoder differ by an order of magnitude in update norm. Single-pool per-layer methods suffer \emph{cross-component budget collapse}, dragging word error rate (WER) far from flat global clipping or collapsing training entirely. When the norm imbalance is milder, adaptive single-pool methods partially recover, confirming that collapse severity scales with the inter-component norm ratio. We empirically diagnose the root cause across six per-layer methods and three speech-LLM architectures. We then propose \emph{$α$-split}, a two-pool allocation that normalises encoder and LLM parameters into independent pools, and show that joint $\ell_2$ sensitivity and the original $(\varepsilon,δ)$-DP guarantee are unchanged. At architecture-calibrated $α$, our method recovers WER utility compared to flat DP, while granting the encoder $4.47{\times}$ tighter per-component noise protection against speaker voice-based gradient-inversion attacks at only $+2.6\%$ LLM noise overhead.

Thu 10 SeptComputation and Language
The gist
Training large speech models with user data needs to keep privacy protected. Existing privacy techniques that treat all parts of the model equally do not work well because speech models have very different behaviors between the sound-processing part and the language part. The authors found that splitting privacy protections into two separate groups for these parts keeps accuracy high and still offers strong privacy guarantees. Their method reduces errors in speech recognition while protecting against attacks that try to recover a person's voice from training data.
Open 2609.11762v1

Segment evidence aware system improves multilingual speech question accuracy

SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ

Abstract: This paper describes our system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge. We adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline. A language model converts timestamped ASR into coherent event spans, which are expanded by a boundary margin and cropped from the original recording. We then synthesize complementary semantic MCQs with Qwen3.6-27B and acoustic MCQs with Gemini~3.1 Flash-Lite, followed by structural, grounding, answer-consistency, and target-model trainability checks, yielding 359,825 verified MCQs across 21 language and accent variants. A text-only probe partitions the data into weak, text-answerable items used for supervised fine-tuning and strong, audio-dependent items used for reinforcement learning with Group Sequence Policy Optimization (GSPO), stabilized by debiased advantages, sequence-level importance correction, and dynamic filtering. Our system obtains 90.92% accuracy on the final official evaluation set.

Thu 10 SeptComputation and LanguageSound
The gist
Understanding questions about conversations in many languages is difficult, especially when answers depend on audio details. The authors developed a system that breaks audio into meaningful parts and creates multiple-choice questions (MCQs) about both text and sound. They train their model first on easier text-based questions, then on harder audio-based ones, improving its ability to handle difficult audio tasks. This approach achieved nearly 91% accuracy on a multilingual speech test.
Open 2609.11355v1

Acoustic and prosodic cues improve speech turn end detection accuracy

Less can be More: What Aspects of Speech Drive End-of-Turn Detection

Abstract: In conversational AI, detecting when a speaker has finished talking is crucial for natural turn taking. While recent work incorporates semantics, the relative contribution of different modalities remains unclear. We present a controlled ablation of acoustic, prosodic, and semantic signals for streaming end of turn detection using a lightweight trimodal classifier. Under identical training conditions, the acoustic prosodic combination achieves the best balance of accuracy and latency, achieving utterance F1 of 0.93 with 7.8% false alarms at 400ms median latency. Adding text increases premature detections without improving performance. Feature space analysis confirms that prosodic features have the strongest class separability, while text representations overlap substantially. These findings suggest that turn-taking is primarily conveyed through intonation and silence patterns rather than semantic completeness, enabling faster and more reliable systems without expensive text inference.

Thu 10 SeptArtificial IntelligenceSound
The gist
Knowing when someone finishes talking is important for smooth conversations with AI. The paper shows that listening to sound patterns and intonation helps computers better guess when a person stops speaking. Surprisingly, understanding the words doesn’t improve this guess and may actually cause errors. The authors found that focusing on how something is said, rather than what is said, leads to faster and more reliable detection of turns in conversation.
Open 2609.11066v1

Semantic uncertainty improves timing predictions in spoken turn-taking

Using Semantic Uncertainty to Estimate Transition Relevance in Turn-taking

Abstract: Turn-taking is a fundamental mechanism that governs when interlocutors speak and listen. Although Spoken Dialogue Systems (SDS) exploit a range of linguistic, acoustic, and non-verbal cues, they produce ill-timed responses in unscripted interaction. A central challenge is anticipating Transition Relevance Places (TRPs), or opportunities, not obligations, for a listener to take the floor. Human listeners do not wait for turn endings; as an utterance unfolds, they use expectations about its developing meaning to anticipate TRPs and decide whether to take the floor. We examine whether these evolving expectations can be modeled through semantic uncertainty -- an LLM-derived measure of how strongly a turn so far constrains what may plausibly come next. To do so, we sample possible continuations of an ongoing turn and use changes in semantic dispersion to identify TRPs within turns. We evaluate this account on a dataset with TRP labels derived from real-time listener responses, rather than retrospective annotation. Our approach substantially outperforms prompt-based and fine-tuned text-only baselines, providing empirical support for the view that evolving semantic constraints inform perceived turn-taking opportunities in unscripted interaction.

Thu 10 SeptComputation and Language
The gist
The problem is figuring out when someone should start talking in a conversation without waiting for the other person to finish completely. The authors show that by measuring how uncertain a computer is about what word might come next in a sentence, it can better guess natural points for switching speakers. They tested this idea using real-time listener reactions and found it outperformed other computer-only methods. This means computers can better predict when to jump into a conversation like humans do.
Open 2609.10934v1

Greek text to speech system improves with small curated audiobook data

Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS

Abstract: Modern TTS systems approach human quality for high-resource languages but degrade when clean speech data is scarce. Modern Greek exemplifies this, lacking the curated corpora behind state-of-the-art synthesis. We propose a data curation recipe that transforms audiobook recordings into TTS-ready data via WhisperX alignment and filtering. Then we fine-tune Parler-TTS (880M), a prompt-based multilingual model whose pre-training encodes phonetic priors transferable to Greek. During development, we find that LLM-generated style prompts introduce speaker drift at inference. Replacing them with deterministic prompts resolves this, and a speaker-specific LoRA stage trained on 3.5 h of single-speaker data anchors identity while updating ~5% of parameters. Our system achieves WER 10.7% (2.9 above the ASR floor), MOS-I 4.00 (vs. 4.36 human speech), and near-human speaker consistency (MOS-C 4.24 vs. 4.30), showing that robust single-speaker Greek TTS is achievable with limited curated data.

Wed 9 SeptSoundComputation and LanguageMachine Learning
The gist
Making computers speak in Greek sounds nearly as good as a human even when there's only a little clean speech to learn from. The authors turned audiobook recordings into good quality training data using special alignment tools. They adapted a large speech model trained on multiple languages to Greek, fixing issues where the voice would change unexpectedly by using consistent prompts and a fine-tuning step that keeps the speaker’s identity. Their system speaks clearly and naturally, close to human voices, using only a few hours of single-speaker recordings.
Open 2609.10022v1

NVV-Locator detects laughter sighs and coughs precisely in speech

NVV-Locator: From Transcript Tags to Acoustic Boundaries for Fine-Grained Nonverbal Vocalization Grounding

Abstract: Human speech includes nonverbal vocalizations (NVVs), such as laughter, sighs, breaths, and coughs, which convey affective and interactional information. Existing approaches typically represent NVVs as transcript-level tags, providing limited supervision for their waveform-time boundaries. We present NVV-Locator for fine-grained NVV temporal grounding. We first unify 26 NVV categories across public resources and construct large-scale timestamp-supervised training data through dual-LLM verification, transcript-guided forced alignment, and energy-based boundary refinement. We further introduce NVV-TimeBench, an expert-refined benchmark with 667 utterances and 1,094 events. NVV-Locator uses a non-autoregressive slot-filling architecture to jointly predict lexical timestamps, NVV categories, and event boundaries. On NVV-TimeBench, it achieves 71.0% Micro F1, 70.2% Macro F1, 80.4% Macro mIoU, and 59.6 ms Macro mMAE, outperforming the evaluated large audio model counterparts. Evaluation on an external corpus further demonstrates the cross-corpus generalization of NVV-Locator.

Wed 9 SeptSound
The gist
People often add sounds like laughter, sighs, and coughs when they talk, and these sounds help express feelings or reactions. The authors developed a system called NVV-Locator that finds exactly when these sounds happen in speech recordings. They combined information from many sources to train their system using precise timing data and created a benchmark set to test it. NVV-Locator outperforms current audio models at detecting these nonverbal sounds and works well even on data it wasn't trained on.
Open 2609.09940v1

Bilingual speech system recognizes words and vocal sounds together

Source-Adaptive Data Curation for Bilingual NVV-Aware ASR

Abstract: Nonverbal vocalizations (NVVs), such as laughter, sighs, breaths, and coughs, convey affective and interactional information that conventional automatic speech recognition (ASR) systems often discard. We present a bilingual Mandarin-English system for Track 1 of the NVVSpeech Challenge at ISCSLP 2026, which requires joint transcription of lexical content and 16 NVV categories at their transcript-relative positions. Our NVV-Aware Whisper adapts Whisper-medium through checkpoint-compatible vocabulary remapping, enabling lexical tokens and inline NVV tags to be decoded within a unified autoregressive sequence without expanding the vocabulary. To provide reliable and diverse supervision, we further introduce a source-adaptive data curation strategy that refines public NVV corpora through acoustic augmentation and multimodal LLM filtering, while mining spontaneous NVVs from in-the-wild media through automated preprocessing and annotation. Under the official bilingual evaluation protocol, the proposed system improves final score from 33.32 to 53.61, with ablations confirming the complementary benefits of the proposed data-curation components.

Wed 9 SeptSound
The gist
Many speech recognition systems ignore sounds like laughter or coughs that people make while talking. The authors created a program that listens to both Mandarin and English speech and also detects 16 types of these nonverbal sounds in the exact spots they happen. They improved an existing speech model so it could handle these extra sounds without making the system more complex. To train it well, they gathered and cleaned up lots of real-world speech data that included these sounds. Their final system works much better than before in recognizing speech and nonverbal sounds together.
Open 2609.09929v1

Speech models struggle with technical talk in science fields

$S^3$-Bench: Evaluating Speech Interaction Models as Scientific Voice Assistants

Abstract: The advance of multimodal large language models (MLLMs) has fundamentally reshaped the paradigm of human-computer interaction, especially speech interaction models capable of seamless conversations. Despite remarkable performance as general voice assistants, their performance in specialized domains remains underexplored, particularly in scientific areas. Scientific interactions introduce formidable challenges, involving rare technical terminology, spoken norms of abbreviations, and the natural verbalization of symbolic special expressions. In this paper, we introduce S$^3$-Bench, a systematic evaluation framework covering 10 major disciplines, consisting of a Knowledge set for speech question-answering and a Dialogue set for multi-turn progressive interactions with simulated user agents. By decomposing a complete atomic turn into stages of speech recognition, perception, knowledge utilization with reasoning, and response pronunciation, we systematically characterize the common challenges and performance tradeoffs of existing approaches. Furthermore, experiments on multi-turn interactions reveal persistent limitations in user adaptation and the generation of accurate, comprehensive, and efficient responses.

Wed 9 SeptComputation and Language
The gist
Voice assistants are good at general conversations but have trouble in scientific areas where people use complex terms and symbols. The authors created S3-Bench, a way to test how well these voice assistants handle science topics across 10 different fields. They look at speech recognition, understanding, reasoning, and speaking responses, finding that current models often fail to adapt well or give complete and accurate answers during longer talks. This shows that more work is needed to make voice assistants better for specialized scientific use.
Open 2609.09852v1

Streaming speech tokenization method cuts latency and improves accuracy

StreamAlign: Streaming Text-Aligned Speech Tokenization

Abstract: Text-aligned speech tokenization methods have emerged to better align speech tokens with LLM token spaces, enabling more effective utilization of pretrained LLMs. However, they rely on offline automatic speech recognition (ASR), leading to two key limitations: (i) the need for complete utterances before tokenization, precluding real-time streaming, and (ii) vocabulary mismatch between ASR and LLMs, which reduces acoustic granularity from the subword to the word level. We introduce StreamAlign, a text-aligned speech tokenization framework that enables streaming tokenization for real-time speech-text joint modeling. StreamAlign performs online speech-text alignment by combining character-level RNN-Transducer alignment with word-level ASR guidance, mitigating ASR-LLM vocabulary mismatch while preserving recognition accuracy. A proactive word boundary classifier anticipates word completion at chunk boundaries, reducing tokenization latency from 560 ms to 270 ms. On LibriSpeech, StreamAlign achieves the lowest WER and highest UTMOS among evaluated tokenizers. Furthermore, StreamAlign-SLM, a spoken language model trained on StreamAlign units, outperforms other end-to-end spoken language models in speech continuation while achieving the strongest overall consistency on SALMon and spoken StoryCloze.

Wed 9 SeptComputation and LanguageSound
The gist
Many systems need to convert spoken words into text tokens before further processing. Existing methods usually wait until you finish speaking before they start, making real-time use impossible. The authors introduce StreamAlign, a new method that aligns speech with text as it happens, predicting word boundaries early to reduce delay. This improves how quickly and accurately speech is turned into tokens, helping machines understand spoken language better and faster.
Open 2609.09719v1

X2-NativeCursor improves real-time text progress tracking in streaming speech

X2-NativeCursor: Native-Token Text Progress Tracking for Incremental-Text Streaming Codec TTS

Abstract: Incremental-text streaming text-to-speech (TTS) needs online text progress tracking for synchronized highlighting, interruption handling, and dialogue-history updates. Input text arrives before it is spoken, so text arrival alone cannot indicate speech progress. Existing waveform-based alignment requires complete audio or adds acoustic processing during streaming. We propose X2-NativeCursor, a lightweight observer that tracks progress from native speech tokens before waveform decoding without changing the TTS generator. Its normalization plan links spoken labels to their original-text spans. Text and native-token encoders feed a local matcher that estimates the current label position. A separate output rule converts revisable position estimates into a cursor that never moves backward. Mean absolute error against an automatic reference is 0.151 Chinese characters with 80-ms lookahead, versus 1.253 characters with 320-ms lookahead for an online waveform baseline. Alignment real-time factor also decreases from 0.3598 to 0.0180 relative to this baseline. Lower tracking error is retained under a second automatic alignment reference. We evaluate X2-NativeCursor on Qwen3-TTS and validate its adaptation to CosyVoice2 by training a separate observer for each backbone. Code is publicly available at https://github.com/X-Square-Robot/X2Streaming-TTS.

Wed 9 SeptComputation and Language
The gist
It’s hard for speech software to keep track of which word is being spoken in real time, especially when the text arrives before the speech is produced. The authors created X2-NativeCursor, a lightweight tool that follows the speech progress by looking at internal speech tokens instead of waiting for audio output. This method tracks spoken words more accurately and faster without changing the main speech generator. It works well for Chinese and adapts to different speech models.
Open 2609.09677v1

Full duplex speech models vulnerable to spoken interruption attacks

DuplexJail: Safety Alignment Breaks Under Spoken Interruption in Full-Duplex Models

Abstract: Full-duplex speech models accept user speech while generating responses, creating an underexplored attack surface. We introduce DuplexJail, which delivers fixed, request-independent spoken prompts through the user audio channel. We compare fixed-delay interruption after the harmful request ends with refusal-triggered interruption following a cue in the model's streaming text. Across four open-source models and 720 harmful requests from AdvBench and HarmBench, fixed-delay interruption raises whole-response attack success rates on AdvBench to 40.3% for PersonaPlex and 48.7% for PersonaPlex-RL, increases of +33.8 and +39.3 percentage points. The refusal-triggered policy reaches 35.6% and 48.6%, respectively, with all trials scored regardless of whether an interruption occurs. Selected conditions also increase FLM-Audio's harmful-response rate, while BayLing-Duplex shows decreases. These findings identify spoken interruption as a jailbreak attack vector and motivate evaluating safety throughout ongoing full-duplex interaction.

Tue 8 SeptCryptography and SecuritySound
The gist
Full-duplex speech models can listen to users while speaking back, but this creates new ways to trick them. The authors found that playing certain fixed spoken messages during or after harmful requests can make these models respond incorrectly, bypassing safety rules. They tested different interruption timings and multiple speech models, showing some models became more likely to produce harmful content. This reveals a new type of security risk for voice assistants and similar systems.
Open 2609.09420v1

TASTE2 enables real-time full-duplex voice interaction with interruption handling

TASTE2: Text-Aligned Speech Modeling and Deployment toward Full-Duplex Voice Interaction

Abstract: Full-duplex voice interaction requires more than utterance-level conversion. It must process streaming speech, manage turn-taking and interruptions, while preserving pretrained linguistic competence and acoustic paralinguistic cues. We ask whether TASTE (Text-Aligned Speech Tokenization and Embedding) provides a viable path toward this goal. We present TASTE2, which transforms utterance-level TASTE into an incremental dialogue stack. A shared text-token vocabulary removes word-level averaging, while modality-aligned dialogue training predicts one continuous audio latent per text token without interleaving heterogeneous token streams. An incremental Speech Detokenizer enables streaming synthesis through CosyVoice2. After speech and dialogue training, TASTE2 (Merge) reaches 56.3% on LLaMA-Questions against a 57.3% Qwen2.5-7B Instruct text-only reference (98.2% accuracy retention), and TASTE2 (Direct) reaches 53.0% (92.4% retention). We build TASTE2 VoiceBot, which processes user speech incrementally, streams synthesized audio, and stops generation on barge-in. On Full-Duplex-Bench v1.0, TASTE2 and TASTE2 VoiceBot handle interruptions well while maintaining high conversational coherence. Natural conversation remains challenging, and deployed mean time to first audio is 2.701 s on two NVIDIA RTX A6000 after TensorRT acceleration. Finally, to our knowledge, we provide the first systematic characterization of explicit paralinguistic control in a TASTE based model. Fast speaking rate serves as a cross-strategy proof of concept after dialogue SFT, while emotion control is strategy dependent and the remaining attributes stay weak. Together, these results establish TASTE based modeling as a practical route toward full-duplex systems while identifying natural conversation robustness, speech generation latency, and feature general paralinguistic control as open challenges. Explore TASTE2 online.

Tue 8 SeptSound
The gist
Having smooth back-and-forth voice conversations with computers is hard because it requires the system to listen and talk at the same time while understanding speech and responding quickly. The authors tested a method called TASTE2 that turns speech into shared text tokens to better handle interruptions and keep conversations flowing naturally. They built a voice robot that can listen and respond in real time, even if the user talks over it. While it works well, there is still room to improve the speed and naturalness of the generated speech.
Open 2609.08956v1

Open source speech model handles talking and editing by instructions

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

Abstract: We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances and 1.95 million hours of effective supervision across five task families: speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing. AuK combines a multimodal large language model for semantic conditioning, an VAE jointly trained on speech, general audio, and music for acoustic conditioning, and a hybrid rectified-flow Transformer that performs dual-stream MMDiT blocks followed by unified single-stream DiT blocks for generation. Training begins with generation-only warm-up and proceeds to joint generation--editing pre-training. We then apply complementary post-training strategies: human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation. To reduce inference cost, we further distill the model with consistency initialization and task-routed Decoupled DMD. The resulting AuK-Flash performs 4-step inference without classifier-free guidance and achieves a 4.5 wall-clock speedup over the full model under matched conditions. Experiments demonstrate leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing, while remaining competitive on signal-level restoration tasks. We release both the source code and model weights to support reproducibility and further research.

Tue 8 SeptSoundComputation and LanguageMultimedia
The gist
Speech technology usually separates creating voices from editing them, but this work combines both in one model that listens to natural language instructions and audio examples. The authors built a huge training set covering many speech tasks and designed a model that mixes understanding meaning and sound together. They also use tricks like human feedback and distillation to make the model faster and better. The final open-source model performs well at generating new speech and editing existing clips just by following spoken or written directions.
Open 2609.08936v1

Speech deepfake detection improved by preserving decision evidence

From Scores to Evidence: Auditable Decisions Can Improve Speech Deepfake Detection

Abstract: Speech deepfakes can mimic a speaker's voice convincingly enough to deceive listeners and automated systems. This has driven strong progress in speech deepfake detection, but most detectors still end with one score per utterance. That score is useful for ranking systems, yet it says little about why a borderline item should be trusted, deferred, or reviewed. Two utterances can fall in the same score band for different reasons, for example because passive and retrieval evidence disagree or because the keyed probe is unavailable. We ask whether the final decision can remain scalar without discarding that provenance. We answer this question with an auditable decision record that carries four aligned cues into a late calibration step: a passive detector score, a conditional keyed-probe score on a marked derivative, retrieval support, and a speaker-profile margin, together with explicit disagreement coordinates. On the 4,080-example ASVspoof 5 Track 1 matched subset, the fixed retrieval-augmented rule improves on retrieval-only evidence, from 15.84 percent to 11.91 percent EER, and late calibration over the full record reaches 8.43 percent EER. At a 33.75 percent review budget, the exposed cue union covers 82.85 percent of the calibrated model's errors. The best passive WavLM run still reaches 6.71 percent EER, so we do not present the decision record as a stronger standalone detector. Its contribution is to preserve the evidence behind each surfaced utterance while still producing one operating score for thresholding and review.

Tue 8 SeptSoundComputation and Language
The gist
Speech deepfakes can trick people and machines by sounding like someone else. Most detection systems just give a simple score for each audio clip, which doesn’t explain why the decision was made or if it should be double-checked. The authors developed a method that keeps detailed clues from different checks and combines them into one score, so it’s easier to know when to trust, review, or reject a speech sample. This method improved detection accuracy and helps reviewers focus on the riskiest cases by showing why the system was unsure.
Open 2609.08899v1

TontaubeV1 enables natural streaming text to speech on consumer GPUs

TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context

Abstract: Text-to-speech systems often face a trade-off between natural prosody and efficient inference: higher perceptual quality typically comes at increased computational cost and latency. We present TontaubeV1, a model that preserves natural prosody while enabling streaming from a single consumer GPU. Speech is encoded by the hierarchical DualCodec representation at 12.5 Hz, which separates a semantic stream from successive acoustic refinements. Our design assumes that prosodic structure is largely established when the semantic stream is generated, and allocates capacity accordingly: a Qwen3-1.7B-derived transformer predicts that stream and thereby the utterance duration, while three progressively smaller Qwen3-0.6B-derived transformers each add one acoustic refinement. Text is tokenized per character rather than by subword. Paired text and audio markers at shared positions support long-form generation with bounded context, and overlapping DualCodec reconstructions are mapped into the VibeVoice acoustic latent space and decoded causally, enabling streaming despite DualCodec's noncausal decoder. The model accepts up to one minute of reference audio for voice conditioning and is designed primarily for English and German, with additional multilingual support. The four predictors total 2.9B parameters; on a single RTX 5090 the streaming path reaches approximately 200 ms to first audio. In separate non-streaming measurements, the end-to-end real-time factor (RTF) is 0.08 for one input and the aggregate RTF is 0.02 across eight concurrent inputs. On our LLM-as-a-judge audiobook-reading benchmark, TontaubeV1 matches ElevenLabs Flash v2.5 and outperforms Fish Audio S2 Pro, the April 2026 Gradium API, and Cartesia Sonic 3 on prosody. The model weights are released on Hugging Face under the Tontaube Community Model License 1.0.

Tue 8 SeptSoundComputation and LanguageMachine Learning
The gist
Making computer voices sound natural while speaking quickly is a challenge because better quality usually means slower processing. The authors present TontaubeV1, a new model that can produce natural-sounding speech in real time using just one standard graphics card. It works by first predicting the basic meaning and timing of spoken words, then adding layers of sound detail step-by-step. This approach lets TontaubeV1 start speaking within 200 milliseconds and supports voices in English, German, and some other languages.
Open 2609.08703v1

X2Streaming-ASR cut delay drastically in live speech recognition

X2Streaming-ASR: wait when uncertain, emit when ready for streaming ASR

Abstract: Streaming automatic speech recognition (ASR) for real-time voice agents and full-duplex dialogue must provide accurate partial transcripts with low commit latency. Existing systems commonly use a fixed chunk size, look-ahead, or target delay, or encourage emissions near estimated acoustic boundaries. These approaches do not directly optimize how much additional context to use at each output position under a single-pass, hard-commit constraint. We propose X2Streaming-ASR, which decomposes streaming recognition into when to commit and what to commit. Its three-stage training procedure first establishes streaming recognition ability, then warm-starts the commit policy with automatically probed trajectories, and finally refines the policy using character-level, segment-assigned group-relative rewards for recognition accuracy and latency. Across AISHELL-1/2/3 and WenetSpeech, X2Streaming-ASR achieves a mean character-level commit latency of 27-84 ms relative to forced-aligned character endpoints, compared with 409-585 ms for the evaluated streaming baselines. It achieves the best streaming CER among the evaluated systems on AISHELL-1 and AISHELL-3 with substantially lower latency.

Tue 8 SeptSoundArtificial Intelligence
The gist
Real-time speech recognition systems often struggle to balance quick response times with accurate partial transcripts. The authors propose a new method called X2Streaming-ASR that decides when to commit to recognizing spoken words and what to commit in a more flexible way. Their approach reduces the delay between hearing and transcribing speech from hundreds of milliseconds to just tens, while improving accuracy on challenging datasets. This could help voice assistants and dialogue systems respond faster and more accurately.
Open 2609.08672v1

Audio visual system improves dialogue clarity in noisy speech

Noise Adaptive Streaming Audio-Visual Speech Token Enhancement for Robust Full-Duplex Spoken Dialogue Models

Abstract: Full-duplex spoken dialogue systems enable simultaneous listening and speaking, but their audio-only perception often fails under background noise and overlapping speech, leading to incoherent responses. Recent audio-visual dialogue approaches show that incorporating visual cues such as lip movements improve robustness under audio corruption. However, existing approaches often adapt the large speech dialogue model itself to process visual input, requiring costly multimodal training. We propose AV-STE, a modular streaming audio-visual front-end that restores corrupted semantic speech tokens from noisy audio and lip video before they reach the speech LLM. The downstream dialogue model remains entirely frozen, preserving its pretrained conversational capabilities. When integrated with frozen Moshi, AV-STE improves average GPT-4o-judged response coherence from 1.42 to 1.91 under same-dataset speaker interference while largely preserving turn-taking behavior. Gains also transfer to out-of-domain Seamless Interaction.

Tue 8 SeptSoundArtificial IntelligenceHuman-Computer Interaction
The gist
Spoken dialogue systems that listen and talk at the same time often get confused when there is background noise or multiple people talking. To fix this, the authors created a special tool called AV-STE that uses both sound and lip reading from video to clean up what is being said before sending it to the conversation brain. This way, the main dialogue system stays the same and still understands and talks well. Their method made the conversation clearer and easier to follow when people were talking over each other or in noisy places.
Open 2609.08390v1

Full duplex speech data created from real conversations using reconstruction and expansion

ConversationalVoice: Full-Duplex Speech Data from Real Conversations through Source-Faithful Reconstruction and Conversation-Grounded Expansion

Abstract: Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backchannel behavior, yet these signals are entangled across speakers in noisy real-world recordings. We present Conversational Voice, a pipeline that converts real two-speaker excerpts into three complementary training-data artifacts. (1) Separation recovers speaker-specific tracks with stable speaker assignments, a canonical transcript, and naturally observed interaction timing. (2) Reconstruction generates speech in matched voices from a fixed source transcript, reconstructs the source turn order, pauses, and overlaps, and adds word-level alignment and delivery instructions. (3) Expansion generates new dialogue constrained by the source context, speakers, and observed interaction pattern. Automatic speaker-verification metrics remain strong across stages, with same-speaker similarity of 0.983-0.991 and positive discrimination margins of 0.199-0.209. Predicted speech quality (NISQA MOS) is 3.56 for separation, 4.41 for reconstruction, and 4.61 for expansion. A Gemini-based automatic evaluator assigns expansion mean scores of 4.94/5 for contextual coherence and 4.80/5 for dialogue naturalness. Expansion and reconstruction exhibit broadly similar interaction profiles; expansion's turn, overlap-event, backchannel, and interruption rates are 4.6%, 8.0%, 13.2%, and 16.0% lower, respectively. We evaluate data properties only; downstream gains in full-duplex model training remain for future work.

Tue 8 SeptComputation and LanguageSound
The gist
Getting computers to handle conversations with natural features like overlapping speech and interruptions requires special training data. This paper introduces a method called ConversationalVoice that starts with recordings of real two-person conversations and creates cleaned-up, speaker-separated audio, rebuilt dialogue that keeps timing and speaker style, and new conversation examples that fit the original context. The authors show that their approach keeps the voices clear and natural sounding, and the generated conversations remain coherent and similar in interaction style. They focus on producing good training data and leave training full-duplex speech models for future work.
Open 2609.08147v1

Tts system generates overlapping two-speaker dialogue with natural timing

KABURI-TTS: Phoneme-Keyed Activity-conditioned Bi-channel Utterance Rendering for Interaction

Abstract: Realizing full-duplex spoken dialogue requires large amounts of two-channel, one-speaker-per-channel conversational speech data. Although conversational text-to-speech (TTS) engines have been developed, they are not necessarily robust to two-party simultaneous phenomena such as backchannels, interruptions, and overlaps that occur while the interlocutor is speaking. In this work, aiming at conversational speech synthesis that reproduces human-like overlap, we propose KABURI-TTS. KABURI-TTS takes a per-speaker phoneme raster as input and renders the speech of the two speakers on separate channels, conditioned on the per-frame phonemes and the voice activity derived from them. Because the phoneme raster is supplied by a separate module, the proposed method enables controllable generation of one-speaker-per-channel, two-party spoken dialogue. A user evaluation shows that, compared with strong baselines, the proposed method attains higher naturalness at both the utterance and the interaction level. Furthermore, an analysis of voice activity confirms that the proposed method produces more overlap and more frequent turn-taking.

Mon 7 SeptSoundComputation and Language
The gist
Making computer voices that talk like people in natural conversations is hard, especially when two people speak at the same time or interrupt each other. The authors created a system called KABURI-TTS that can generate speech for two speakers separately but in a way that sounds like real overlapping talks. It uses detailed timing of sounds (phonemes) and when each speaker talks to decide how to sound. Tests showed it sounds more natural and captures the way people interrupt or talk over each other better than older methods.
Open 2609.07200v1