OpenEnded speech corpus improves English speaking assessment accuracy
OpenEnded: An Open-Response Speech Corpus for Speaking Proficiency Assessment with Human Annotations and ALM Supervision
Sound
Summary
Speaking assessment tools often struggle because they rely on people reading aloud, which doesn't match real conversations. The authors created OpenEnded, a collection of English speech samples from Mandarin speakers answering open-ended questions, which is closer to how people actually speak. They labeled parts of the speech for accuracy, fluency, and tone, using both human reviewers and an AI model. This dataset helps make automatic speaking assessments better, as shown by testing different models on it.
What this means in practice
- •For language learning app developers: Integrate detailed automated speech scoring for real answers to improve user feedback on pronunciation and fluency.$Commercial implications: Enables enhanced speaking practice and assessment features that can be sold to language learners as part of apps.
- •For assessment platform engineers: Train speech evaluation models on utterance-level annotated open-response data to improve assessment realism and accuracy.
Authors
Yu-Wen Chen, Eric Zhou, Evelyn Ding, Tianyi Shen, Zhou Yu, Julia Hirschberg
Abstract
The development of automated speaking assessment (ASA) is limited by the scarcity of public datasets, with most existing work relying on read-aloud speech, which limits applicability to real-world communication scenarios. In this work, we introduce OpenEnded, a corpus of English practice speech from Mandarin speakers in open-response tasks. Unlike prior open-response datasets that provide only holistic proficiency scores, OpenEnded offers utterance-level assessments of accuracy, fluency, and prosody. Approximately 10,000 utterances are collected and annotated using a hybrid framework: 1,000 are manually labeled via multi-rater scoring with discrepancy resolution to form a high-quality test set, while the remaining are pseudo-labeled by an audio language model (ALM) for training and development sets. We evaluate ALMs and existing ASA models on the OpenEnded test set and introduce VoxPA as an additional baseline. Results show that ALM-generated pseudo-labels improve training over original ALM scoring, while VoxPA achieves the best performance among all baselines.