MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction
2026-08-24 • Artificial Intelligence
Artificial Intelligence
AI summaryⓘ
The authors present MediSkill-Evo, a clinical agent designed to improve medical diagnoses and treatment planning while following proper clinical procedures. Unlike some systems, it organizes experience into four categories and uses safety checks to make reliable decisions without altering its core model. Their tests show that MediSkill-Evo performs better than a previous system called AgentClinic on diagnosis accuracy, treatment coverage, and reducing critical errors. The study focuses on system-level performance under various challenging conditions but does not claim clinical validation or evaluate individual components independently.
clinical agentpartial observabilityprocess knowledgediagnosis accuracytreatment-intent coveragesafety checkssymbolic schemasbackbone modelclinical processevaluation suite
Authors
Ruoyu Wu, Shenfu Xie, Yinqian Sun, Haibo Tong, Feifei Zhao
Abstract
Interactive clinical agents must gather decisive evidence and convert it into grounded actions under partial observability. A correct final diagnosis alone does not show that an agent respected evidence and care-process constraints. We introduce MediSkill-Evo, a clinical agent that evolves governed process knowledge without backbone fine-tuning. It separates experience into four typed banks for clinical skills, process rules, symbolic schemas, and measurement procedures. Provenance, support, replay, and controller-defined safety checks govern publication to a frozen test-time snapshot. A Process-Constrained Preference Harness binds evidence to its source, rejects controller-invalid candidates, and ranks actions with a safety-prioritized Clinical Process Critic. We evaluate complete agent systems across two backbone endpoints and six controlled stress dimensions under the same Doctor-turn limit. On 300 held-out Qwen encounters, MediSkill-Evo improves diagnosis accuracy from 61.33 percent to 69.00 percent and treatment-intent coverage from 33.62 percent to 66.44 percent, while reducing automatically scored critical failures from 31.00 percent to 16.33 percent relative to AgentClinic. On 180 hard-isolation conditions derived from 30 cases, target recovery reaches 93.61 percent under patient-behavior pressure, 100.00 percent for temporal evidence, and 92.22 percent for triage red flags. An exploratory 100-case MedSAM comparison evaluates request-gated tool-interface feasibility. These results provide descriptive end-to-end evidence for the complete system on fixed evaluation suites, not causal evidence for an individual bank or clinical validation of the automatic judge.