Cross-language model improves respiratory disease detection from speech
A Cross-Lingual Acoustic Disease-Alignment Framework for Respiratory Health Assessment from Spontaneous Speech
SoundComputation and Language
Summary
Detecting respiratory diseases from spoken language is tricky because different languages sound very different. The authors created a new method called CL-DAF that finds speech features linked to disease that stay consistent across languages. They tested it on English and Bangla speakers and improved the accuracy of identifying lung disease compared to existing methods. This approach helps build speech-based health tools that work for many languages.
What this means in practice
- •For healthcare technology developers: Build multilingual speech analysis tools to screen respiratory diseases across different language populations.
- •For public health monitoring teams: Use aligned acoustic speech features to monitor respiratory health trends in diverse communities speaking different languages.
Authors
Roksana Khanom, Raghib Asfak Tasnim, Bodrun Nahar Bithi, Shafia Shirin Supty, Saiful Islam Raju, Ashok Agrawala, Nirupam Roy
Abstract
Spontaneous speech offers a scalable, noninvasive signal for respiratory health assessment, yet interpretable models that generalize across languages remain challenging because disease-related acoustic changes are confounded by language-specific phonetic variation. We present CL-DAF, a Cross-Lingual Disease-Alignment Framework that identifies acoustic dimensions whose disease effects remain consistent across languages. Using 201 English and 75 newly collected Bangla speakers, we construct a common 272-dimensional acoustic representation and quantify disease alignment using signed rank-biserial effects and the Language Invariance Score. We first show that spontaneous Bangla speech separates COPD from controls (AUC 0.85); however, 133 features reverse their disease direction across languages and the full representation transfers poorly (AUC 0.49 from Bangla to English). CL-DAF isolates 26 disease-aligned features that raise AUCs to 0.825 and 0.722 from English to Bangla and Bangla to English, respectively. These findings provide a foundation for multilingual clinical speech models emphasizing pathology over language-dependent variation.