Summary
Doctors in India often have very busy and noisy clinics where they speak many languages mixed together during short appointments. Current voice-recognition systems for taking notes were mainly made using data from English-speaking countries and don’t work well in India’s conditions. The authors looked for real-world examples of doctor-patient talks in India but found almost none available to test these systems properly. They also talked with groups building these systems who use their own ways to check how well things work, making it hard to compare or trust them. The authors suggest creating a shared, multilingual set of real conversations to better evaluate these systems for Indian healthcare.
What this means in practice
- •For healthcare technology developers: Develop shared, realistic multilingual benchmarks to assess clinical speech recognition systems in Indian healthcare settings.
- •For health system procurement teams: Use standardized, real-world evaluation data to compare and select ambient clinical scribes suited for Indian clinical environments.
A position paper. It proposes an approach and reports no results.
Authors
Siddharth D Jaiswal, Krithi S, Ashish Makani, Suvrankar Datta, Sunayana Sitaram, Mohit Jain
Abstract
Ambient clinical scribes (ACS) are being rapidly deployed at scale across Global South healthcare settings, aiming to reduce clinician documentation time, especially in overburdened environments like India. These ACS are primarily developed or distilled from models built and validated on Global North speech, languages and consultation styles. Indian clinical encounters are brief, triadic, multilingual, code-mixed with low-resource languages, and conducted in highly resource-constrained, noisy settings -- increasing the likelihood of ASR and note-generation errors manyfold. We posit an urgent need to develop a standardized evaluation infrastructure to assess whether these systems are safe, reliable, and well-suited to the Indian healthcare setting. We substantiate our claims through a mixed-methods study -- a systematic survey of publicly available patient-clinician conversational datasets, a quantitative comparison of these datasets against conversational and cultural markers drawn from the Indian clinical-communication literature, and semi-structured interviews with five organizations building and deploying ACS in India and Africa. Our survey shows that there are no publicly available, large-scale, real-world benchmarks for ACS in India, with existing datasets being overwhelmingly synthetic. We note that the available Global North datasets diverge significantly from the expected conversational and cultural structures of Indian encounters. Finally, our interviews reveal that deploying organizations have each built proprietary, incomparable evaluation pipelines, creating a fragmented ecosystem with no independent and reliable basis for procurement. We call for the development of a publicly shared, real-world, multilingual benchmark for ACS evaluation and outline the properties and policies such a benchmark would require.