Audio language models struggle with reasoning on stuttered child speech
Reasoning Beyond Transcription: Audio Language Models on Child Stuttering Speech
Computation and Language
Summary
Understanding speech from children who stutter is hard for computers because their speech sounds and patterns are different from adults and include repeated words or breaks. The authors tested special audio language models to see how well they could understand and summarize what stuttering children said during conversations with adults. They asked the models to focus only on the child's speech and keep disfluencies that are important for speech therapy. The findings show that these models can get the general meaning but have a harder time reasoning accurately when the speech becomes more disfluent or mixed with adult voices.
Audio Language ModelsChild speechSpeech disfluenciesStutteringSemantic summarizationSpeech entailmentProsodyMixed-speaker settingsInstruction-guided modelsReference-based metrics
Authors
Chibuzor Okocha, Christan Grant, Zoey Liu
Abstract
Child speech differs from adult speech in acoustics, prosody, and linguistic structures. Speech disfluencies (such as repetitions) further challenge automatic understanding. While Audio Language Models (ALMs) show strong semantic reasoning from speech audio, their ability to reason about disfluent child speech in mixed-speaker settings remains unexplored. We investigate this through two tasks: child-focused semantic summarization and speech entailment. Experiments use recordings of children who stutter in mixed speaker interviews without explicit speaker separation. Models are instruction-guided to focus on the child, preserve clinically relevant disfluencies, and avoid adult-speech leakage. Evaluation combines LLM-based judges and reference-based metrics, anchored by transcript-oracle baselines to isolate errors. Results show that while ALMs extract high-level meaning from stuttered speech, reasoning degrades significantly with increased