Multimodal ai screens autism early using video audio and dialogue
A multimodal large language model for evidence-based autism spectrum disorder screening
Computer Vision and Pattern RecognitionHuman-Computer InteractionMachine Learning
Summary
Early screening for autism is hard because specialists are limited and usual tests can be subjective. The authors created ASDchat, an AI model that looks at videos, audio, and conversations to spot signs of autism and explain its decisions based on clinical criteria. It tested well on over 1,000 kids and worked on new data from different locations. It also grouped autism cases into subtypes and suggested targeted interventions. This tool could help doctors screen more children accurately and consistently.
What this means in practice
- •For clinical practitioners: Provide early and explainable autism screening by analyzing video, audio, and dialogue data with aligned clinical criteria.
- •For healthcare technology developers: Integrate multimodal AI screening tools into digital health platforms to enhance autism detection and subtype identification.$Commercial implications: Enables creation of AI-powered autism screening software sold to hospitals and clinics for early diagnosis support.
Authors
Jun Chen, Qi Zhao, Yunliang Jiang, Shuqin Cao, Yunqiang Lin, Chenglong Jia, Qiang Guo, Guang Dai, Xiongtao Zhang, Mengmeng Wang, Xiaoyue Ma
Abstract
The clinical management of autism spectrum disorder (ASD) faces a bottleneck in early screening, mainly because trained specialists are scarce and conventional assessment tools are subjective. Here, we introduce ASDchat, a multimodal large language model designed for evidence-based ASD screening, which takes video, audio, and dialogue as input. ASDchat adopts a dual-branch architecture, where the decision branch generates screening probabilities and the evidence branch generates traceable, timestamped behavioral evidence aligned with standardized clinical criteria (ADOS-2). The model was trained and evaluated on a dataset of 1,035 participants from 27 sites in China, which covered typically developing (TD) children, children with ASD, and children with other disorders. For ASD versus TD, ASDchat reached an area under the receiver operating characteristic curve (AUC) of 0.953 $\pm$ 0.021. On 9 held-out sites that were not used for training, the mean AUC was 0.932. Furthermore, unsupervised clustering of the behavioral dimensions split the ASD cases into six subtypes with different phenotypic profiles, and ASDchat suggests an intervention for each subtype. ASDchat provides a feasible path for large-scale, evidence-based early ASD screening in clinical practice.