Papers for

clinical ai quality assurance teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Hallucinations expose gaps in medical AI evaluation methods

When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation

Abstract: Hallucinations can undermine clinician trust in LLMs, making it important that evaluation methods capture clinically relevant errors. Rubric-based evaluation has become the leading approach for assessing LLMs in medicine, but it is unclear whether rubric scores reflect such errors. We first study this in a controlled setting using MedHallu, finding that more specific rubrics better distinguish correct from hallucinated responses. To test this systematically, we develop a taxonomy of medical hallucination types and a clinician-validated error-injection pipeline that creates matched correct and error-injected responses. Across HealthBench, HealthBench Professional, and LiveMedBench, our clinically relevant hallucinations are missed by rubrics, often leaving scores unchanged. We find that rubrics are most effective when explicitly checking facts, and are less effective for additional or unexpected errors they do not anticipate. A preliminary retrieval-based factuality check recovers some of the rubric-blind errors, suggesting a complementary approach. These findings reveal systematic blind spots in current medical evaluation of LLMs and suggest that rubric scores alone are insufficient to establish clinical reliability, potentially undermining clinician trust and confidence in clinical deployment.

Fri 11 SeptArtificial Intelligence
The gist
Medical AI systems sometimes make up incorrect information, which can reduce doctors' trust in them. The authors found that current evaluation rubrics do not always catch these errors, especially unexpected or unclear ones. They developed a new method to create and classify these mistakes to test how well evaluations work. Their study shows that relying on rubrics alone is not enough to ensure medical AI is trustworthy for clinical use.
Open 2609.12718v1