Multi agent ai improves clinical interview training outcomes

Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training

Multiagent SystemsArtificial IntelligenceHuman-Computer Interaction

Summary

Medical students need practice conducting patient interviews, which is often costly and hard to scale. The authors created an AI system with multiple agents that simulate a patient, tutor, and evaluator to help train students. In a study, this system helped students communicate better and show more empathy without increasing diagnostic mistakes. The authors also shared a detailed dataset to support future research on AI in clinical training.

What this means in practice

  • For medical training centers: Use multi-agent AI systems to provide scalable and effective patient interview practice for medical students.
  • For virtual healthcare simulation providers: Integrate multi-agent language models to enhance clinical communication and empathy training in virtual patient simulations.$Commercial implications: Enables development of advanced AI-driven clinical training products with improved interpersonal skills focus.

Authors

Luming Yang, Haoxian Liu, Siqing Li, Rong Jia, Yue Xiao, Guanhua Chen, Li Lu

Abstract

Clinical education must prepare medical students to conduct safe and coherent patient interviews under conditions of uncertainty. Traditional standardized patient (SP) training is resource-intensive and difficult to scale. We developed a scaffolding-oriented multi-agent Large Language Model (LLM) AI Standardized Patient (AI-SP) training platform1. The system includes a patient agent for simulated dialog, a tutor agent providing Socratic prompts without disclosing diagnostic information, and a turn-level evaluator agent that monitors clinical progress without revealing summative scores. In a randomized controlled study (N = 100 medical students), participants were assigned to either a multi-agent (MA) scaffolding condition or a control condition. All students completed two learning sessions under their assigned condition followed by an examination conducted in a patient only environment. Performance was assessed using a standardized Objective Structured Clinical Examination (OSCE) based rubric. While no significant difference was observed in final diagnostic accuracy between groups, the multi-agent AI standardized patient system improved final examination scores compared to the control group utilizing structured progressive information disclosure; the most substantial and consistent improvements were observed in communication, the expression of empathy, and specific history-taking behaviors. These findings suggest that specialized LLM agents enhance the process quality of simulated clinical interviews without artificially inflating examination outcomes. To support future research, we release a multi-expert annotated dataset comprising transcripts, checklist annotations, turn-level evaluations, and OSCE-aligned scoring outcomes. This resource aims to facilitate the development of pedagogically grounded AI-SP systems and advance research on AI-supported clinical reasoning training.