Papers for

healthcare technology developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Transformer combines multiple camera views for better 3D human pose estimates

STA-TFM: Spatio-Temporal Aggregation Across Views TransForMer for Pose Estimation

Abstract: Monocular 3D human pose estimation (HPE) remains challenging due to depth ambiguity, occlu- sions, and the need for temporal consistency. While multi-view methods provide superior accuracy over monocular approaches, they often require complex setups. We introduce STA-TFM, a transformer-based architecture that combines spatial and temporal information for multi-view pose estimation. The approach leverages DSTformer, a monocular feature extractor, to capture long-range pose dependencies within each view. A fusion transformer then aggregates information across views to produce coherent 3D estimates. To address training data scarcity, we use a data generation pipeline that transforms any existing 3D pose dataset into multi-view setups with controllable parameters. Experiments on various datasets demonstrate that STA-TFM outperforms existing camera-parameter-free multi-view methods. STA-TFM achieves 50.9% and 49.5% reductions in mean per joint position error (MPJPE) and mean per joint velocity error (MPJVE) on the DHP19 dataset. Furthermore, it achieves 6.7% and 7.7% respective reductions on HAA4D, and a 15.2% MPJPE reduction on TotalCapture. STA-TFM handles noisy and missing 2D inputs, supporting potential deployment in healthcare monitoring, athletic assessment, and immersive technologies. Code, training checkpoints, and data are available at https://zenodo.org/records/22832620.

Mon 21 SeptComputer Vision and Pattern Recognition
The gist
Estimating 3D human poses from a single camera is hard because the camera can only see from one angle, making it tricky to understand depth and handle blocked views. The authors created a new method called STA-TFM that uses a transformer model to combine information over time from several cameras to make more accurate 3D pose estimates. They also developed a way to create training data with different camera setups to help the model learn. Their method works better than other approaches that don’t rely on camera settings and handles missing or noisy input well.
Open 2609.24482v1

Bedside robot asks questions before alerting for distress cues

Ask Before It Tells: Benchmark-to-Robot Body-Cue Transfer for a Question-First Bedside Robot

Abstract: Body-cue recognition can support assistive robots, but benchmark accuracy does not guarantee reliable behavior under a robot-camera viewpoint. We present Nuni, a bedside robot prototype that treats a detected distress cue as a reason to ask rather than a reason to alert. We compare two X3D-UGT RGB appearance classifiers, which reach 97.7% and 94.8% six-way accuracy on NTU RGB+D, with a pose-centric hybrid pipeline on 28 single-actor scripted clips recorded from the robot camera. The hybrid path achieved 0.71 six-way macro recall, versus 0.25 and 0.29 for the fine-tuned and from-scratch RGB variants. More importantly for interaction, it produced a question-triggering distress cue in 12/16 distress clips and would have prompted unnecessarily in 2/8 normal clips; the RGB variants yielded a question-triggering cue in only 2/16 and 3/16 distress clips. We separately tested the question-first controller through event injection. All 13 state-transition trials passed: valid responses caused stand-down, two unanswered prompts produced one alert, and three boundary conditions were handled correctly. These results are a preliminary technical evaluation, not a user study or medical validation, but they show how interaction policy can limit the consequences of uncertain perception.

Mon 21 SeptRoboticsHuman-Computer Interaction
The gist
Recognizing signs of distress is important for assistive robots, but what works well in lab tests may not work well from a robot’s camera view. The authors built a bedside robot called Nuni that asks questions when it sees a possible distress signal instead of immediately raising an alarm. They showed that a hybrid method focusing on body pose recognized distress better from the robot’s viewpoint than appearance-based methods. By asking first, the robot can avoid false alarms and only alert when needed. This approach helps manage uncertainty in robot perception during interactions.
Open 2609.24099v1

Cross-language model improves respiratory disease detection from speech

A Cross-Lingual Acoustic Disease-Alignment Framework for Respiratory Health Assessment from Spontaneous Speech

Abstract: Spontaneous speech offers a scalable, noninvasive signal for respiratory health assessment, yet interpretable models that generalize across languages remain challenging because disease-related acoustic changes are confounded by language-specific phonetic variation. We present CL-DAF, a Cross-Lingual Disease-Alignment Framework that identifies acoustic dimensions whose disease effects remain consistent across languages. Using 201 English and 75 newly collected Bangla speakers, we construct a common 272-dimensional acoustic representation and quantify disease alignment using signed rank-biserial effects and the Language Invariance Score. We first show that spontaneous Bangla speech separates COPD from controls (AUC 0.85); however, 133 features reverse their disease direction across languages and the full representation transfers poorly (AUC 0.49 from Bangla to English). CL-DAF isolates 26 disease-aligned features that raise AUCs to 0.825 and 0.722 from English to Bangla and Bangla to English, respectively. These findings provide a foundation for multilingual clinical speech models emphasizing pathology over language-dependent variation.

Wed 16 SeptSoundComputation and Language
The gist
Detecting respiratory diseases from spoken language is tricky because different languages sound very different. The authors created a new method called CL-DAF that finds speech features linked to disease that stay consistent across languages. They tested it on English and Bangla speakers and improved the accuracy of identifying lung disease compared to existing methods. This approach helps build speech-based health tools that work for many languages.
Open 2609.19398v1

Ambient clinical scribes need multilingual Indian speech data sets

Evaluating Ambient Clinical Scribes in India: The Need for Multilingual Real-World Clinical Conversation Data

Abstract: Ambient clinical scribes (ACS) are being rapidly deployed at scale across Global South healthcare settings, aiming to reduce clinician documentation time, especially in overburdened environments like India. These ACS are primarily developed or distilled from models built and validated on Global North speech, languages and consultation styles. Indian clinical encounters are brief, triadic, multilingual, code-mixed with low-resource languages, and conducted in highly resource-constrained, noisy settings -- increasing the likelihood of ASR and note-generation errors manyfold. We posit an urgent need to develop a standardized evaluation infrastructure to assess whether these systems are safe, reliable, and well-suited to the Indian healthcare setting. We substantiate our claims through a mixed-methods study -- a systematic survey of publicly available patient-clinician conversational datasets, a quantitative comparison of these datasets against conversational and cultural markers drawn from the Indian clinical-communication literature, and semi-structured interviews with five organizations building and deploying ACS in India and Africa. Our survey shows that there are no publicly available, large-scale, real-world benchmarks for ACS in India, with existing datasets being overwhelmingly synthetic. We note that the available Global North datasets diverge significantly from the expected conversational and cultural structures of Indian encounters. Finally, our interviews reveal that deploying organizations have each built proprietary, incomparable evaluation pipelines, creating a fragmented ecosystem with no independent and reliable basis for procurement. We call for the development of a publicly shared, real-world, multilingual benchmark for ACS evaluation and outline the properties and policies such a benchmark would require.

Tue 15 SeptComputers and SocietyHuman-Computer Interaction
The gist
Doctors in India often have very busy and noisy clinics where they speak many languages mixed together during short appointments. Current voice-recognition systems for taking notes were mainly made using data from English-speaking countries and don’t work well in India’s conditions. The authors looked for real-world examples of doctor-patient talks in India but found almost none available to test these systems properly. They also talked with groups building these systems who use their own ways to check how well things work, making it hard to compare or trust them. The authors suggest creating a shared, multilingual set of real conversations to better evaluate these systems for Indian healthcare.
Open 2609.17355v1

Multimodal ai screens autism early using video audio and dialogue

A multimodal large language model for evidence-based autism spectrum disorder screening

Abstract: The clinical management of autism spectrum disorder (ASD) faces a bottleneck in early screening, mainly because trained specialists are scarce and conventional assessment tools are subjective. Here, we introduce ASDchat, a multimodal large language model designed for evidence-based ASD screening, which takes video, audio, and dialogue as input. ASDchat adopts a dual-branch architecture, where the decision branch generates screening probabilities and the evidence branch generates traceable, timestamped behavioral evidence aligned with standardized clinical criteria (ADOS-2). The model was trained and evaluated on a dataset of 1,035 participants from 27 sites in China, which covered typically developing (TD) children, children with ASD, and children with other disorders. For ASD versus TD, ASDchat reached an area under the receiver operating characteristic curve (AUC) of 0.953 $\pm$ 0.021. On 9 held-out sites that were not used for training, the mean AUC was 0.932. Furthermore, unsupervised clustering of the behavioral dimensions split the ASD cases into six subtypes with different phenotypic profiles, and ASDchat suggests an intervention for each subtype. ASDchat provides a feasible path for large-scale, evidence-based early ASD screening in clinical practice.

Tue 15 SeptComputer Vision and Pattern RecognitionHuman-Computer InteractionMachine Learning
The gist
Early screening for autism is hard because specialists are limited and usual tests can be subjective. The authors created ASDchat, an AI model that looks at videos, audio, and conversations to spot signs of autism and explain its decisions based on clinical criteria. It tested well on over 1,000 kids and worked on new data from different locations. It also grouped autism cases into subtypes and suggested targeted interventions. This tool could help doctors screen more children accurately and consistently.
Open 2609.16464v1

Emergency dispatchers identified key step for helpful AI assistance

Situated Action in Pre-Hospital Critical Care Dispatch: Identifying where and how Algorithmic Assistance might be useful in the daily work of specialist Emergency Medical Dispatchers

Abstract: This study uses ethnographic immersion and observation as contextual inquiry to understand situated action at an Emergency Medical Dispatch critical care hub. The work is a response to the urgent need to recruit context-specific knowledge and participation into design that helps to narrow the AI Chasm - the gap between the promise of Artificial Intelligence (AI) systems and what they deliver for clinicians and their patients. The work of a pre-hospital critical care team's dispatch process is described and analysed to reveal both the structure of the workflow and the different cognitive demands it makes on staff tasked with dispatch decision-making. Elements of attention, communication and focus between the humans, as they carry out this work, are drawn out in order to understand the work-as-done and identify the many dependencies in the process. The motivation is to establish where AI support might be useful and to discover what challenges there could be in designing appropriate algorithmic assistance. We ask whether, where and how the design and implementation of an AI system might be considered. The ultimate objective is to improve the decision process itself to the benefit of clinicians and patients. The study identifies three key steps in the situated workflow and details how decision-makers negotiate each one as emergency calls follow complex routes between them. We find compelling evidence that the first of these decision steps constitutes the most promising candidate for unobtrusive assistance that could be safe and effective in improving both clinician workload and clinical outcomes.

Mon 7 SeptHuman-Computer Interaction
The gist
Emergency dispatchers have a complex job making quick decisions when people call for critical medical help. The authors observed dispatchers closely to understand how they work and where AI might support them safely and effectively. They found that the first step in the dispatch decision process stands out as the best place to provide AI help without disrupting the dispatchers' work. This support could reduce dispatcher workload and improve outcomes for patients.
Open 2609.07705v1