AI summaryⓘ
The authors created a special language model called MIND, trained only on trusted patient education materials, to answer questions about the psychiatric medication escitalopram. They compared MIND’s answers to those from ChatGPT and another tool, checking how accurate, clear, complete, nuanced, safe, and helpful the answers were. Computer scoring showed MIND did best overall, but when psychiatrists rated the answers, they found ChatGPT slightly more accurate and preferred its responses more often, despite MIND giving more complete answers. The study suggests MIND can safely provide good information, but doctors still liked ChatGPT’s style better. This is an early step in making language models that give safe, reliable health info.
large language modelLLMpsychiatric medicationspatient educationescitalopramclinical fidelityaccuracysafetyChatGPTpsychiatrist ratings
Authors
Alexander J. Hish, Arjun Nagendran, Scott N. Compton
Abstract
Background This study was designed to evaluate whether a domain-specific large language model (LLM) trained exclusively on patient education resources can answer questions about psychiatric medications, in a manner superior to LLM chatbots. We developed an LLM ("MIND") fine-tuned for clinical fidelity, trained on patient education resources from authoritative medical organizations. Methods We compared the responses of MIND, ChatGPT, and OpenEvidence to patient questions about escitalopram, using two methods: (1) computer analysis according to a rubric measuring accuracy, clarity, completeness, nuance, safety, and referral appropriateness; (2) ratings from N=10 board-licensed psychiatrists on similar metrics. Results When rated by rubric, MIND was rated highest in all domains (p<0.001). When rated by psychiatrists, ChatGPT was rated accurate more often than MIND with a negligible effect size (p=0.021, r=0.073); MIND was rated complete more often than ChatGPT with a small effect size (p<0.001, r=0.160); and MIND and ChatGPT were rated safe with the same frequency (p=0.955, r=0.002). The majority of psychiatrists preferred the responses generated by ChatGPT (57.6%) compared to MIND (42.4%, p=0.003). Conclusions MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time. However, despite MIND's ability to provide more complete responses, psychiatrists preferred ChatGPT's responses. MIND represents a step towards building safe LLM systems to enhance patient education in psychiatry.