Clinical models struggle to diagnose multiple co-occurring diseases in conversations
CLIMB: A Clinical Multimorbidity Benchmark for Diagnosing Co-occurring Conditions through Multiturn Conversations
Machine LearningArtificial IntelligenceComputation and Language
Summary
Doctors often see patients with more than one illness at the same time, and figuring out all these illnesses takes many questions during a visit. The researchers created CLIMB, a way to test computer programs by simulating doctor-patient conversations to see if the programs can identify all the illnesses correctly. They found that current models usually focus on just one main diagnosis and miss other conditions even when given all the information. Asking more questions often leads to more mistakes rather than better diagnoses. This shows that computer models need improvement to handle complex cases with multiple diseases.
What this means in practice
- •For medical ai developers: Test and improve diagnostic models on complex multimorbidity cases using the CLIMB benchmark with simulated multi-turn patient interactions.
- •For healthcare it teams: Evaluate existing clinical decision support tools for their ability to identify multiple co-occurring conditions during patient intake conversations.
Authors
Yusuf Kesmen, Aniruddha Mukherjee, Yena Chang, David Sasu, Trevor Brokowski, Alexandra V. Kulinkina, Kristina Keitel, Akhil Arora, Lars Henning Klein, Mary-Anne Hartley
Abstract
Patients often have several co-occurring clinical conditions, and the findings needed to identify and disambiguate them emerge over the course of a consultation. Evaluating clinical reasoning in this setting requires both multi-turn interaction and multi-label diagnosis. We introduce CLIMB, a benchmark in which a doctor model interviews a simulated patient to recover a ground truth set of co-occurring clinical conditions. Cases are synthesized from clinical decision algorithms and diagnostic datasets, grounding multimorbid presentations in structured clinical knowledge. Across six frontier and open models, none recovers the exact set of conditions in more than 10% of interactive cases. Diagnostic performance declines when conditions co-occur, even when models receive the full clinical record and the true number of conditions. Interaction reduces performance further. In controlled experiments, models behave like single-hypothesis trackers: they anchor on the diagnosis suggested by the opening findings, keep questioning around it, and recover a second condition mainly when a finding in view points to it. Questioning them further does not complete the set but adds mostly wrong diagnoses. We formalise this pattern with a theoretical reference model of single-hypothesis tracking. The benchmark, generator, and evaluation code are available at https://anonymous.4open.science/r/CLIMB-8340.