Clinical experts struggle to guide large language models in data extraction
"I Know Where to Look," But Does the LLM? Charting the Gaps Between Clinical Expert Needs and Unstructured Data Abstraction Tools
Human-Computer Interaction
Summary
Clinical researchers need to pull important facts from complex patient records to study diseases like cancer. The authors created a tool using large language models to help with this, but found that doctors often have tough times teaching the AI what details matter and how to extract them accurately. Challenges included judging which notes are reliable and tailoring the AI prompts. The study shows that current AI tools don’t fully match what clinical experts need and points out areas that need improvement.
What this means in practice
- •For hospital data teams: Improve clinical record annotation workflows by adapting AI tools to better reflect experts’ judgments on note reliability and complex concepts.
- •For legal document reviewers: Develop language model-based systems that require nuanced human input to extract structured data accurately from lengthy unstructured texts.
Authors
Venkatesh Sivaraman, Rigney Turnham, George Bonano, Nevin Aresh, Renumathy Dhanasekaran, Margaret Guo, Sindhu Kubendran, Olivia Lin, Jonathan D Louie, Kristan Olazo, Jeanne Shen, Harish Vasudevan, Jeanette Wong, Emily Alsentzer, Jason A Fries, Anobel Odisho, John Gordan, Jean Feng, Julian C Hong
Abstract
Clinical data abstraction, the process of distilling structured information from patient records, plays a key role in advancing knowledge about diseases such as cancer. Information extraction (IE) with large language models (LLMs) could accelerate this process, but it is unclear whether current frameworks effectively support clinical researchers without AI expertise. To address this, we co-designed an interactive LLM-based abstraction system called Libretto with seven cancer research teams, then evaluated the system's ability to help them answer real-world research questions. We found that while clinicians knew where and how to annotate complex concepts in patient notes, in twelve of fourteen tasks they faced barriers to replicating those intuitions with LLMs. Contextual note reliability judgments, difficulties in steering vibe-coded prompts, and inflexible evaluation strategies necessitated fundamental changes to the IE workflow. Our results highlight open problems for HCI research to bridge the gaps between AI data work tools and clinical users' needs.