Clinical AI system efficiency measured by effort reduction across specialties
KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI
Artificial Intelligence
Summary
Current ways to measure how well clinical AI systems work focus on how similar their output is to examples, not on how much they actually help doctors save time. The authors created KnowBench, a benchmark that measures how much a clinical AI reduces the work doctors have to do by checking how much of the AI’s work is accepted without changes. They tested it across many clinical tasks and found their AI models reduced doctors’ effort by about 98% overall. This new way to measure AI’s usefulness helps compare different systems fairly based on the effort saved.
What this means in practice
- •For clinical operations teams: Measure and compare how much different AI tools reduce clinician effort in generating visit notes and billing tasks using a standardized, auditable metric.
- •For healthcare software developers: Develop AI products with built-in feedback mechanisms aligned to effort reduction benchmarks for clinical documentation and decision support.$Commercial implications: This paper enables building clinically validated AI software that demonstrably decreases clinician workload, which can be marketed to healthcare providers.
Authors
Jocelyn Kang, Caroline Zhang
Abstract
Clinical AI systems are evaluated with instruments built for research settings (reference-based similarity metrics and expert rubric panels) that measure resemblance to an artifact rather than reduction of a burden. We introduce KnowBench, pioneered by Knowtex, whose unifying metric is Effort Reduction (ER): the proportion of system-generated clinical work product accepted by the responsible clinician under expert and safety review. ER is defined once and instantiated per task across the administrative workload clinical AI automates: visit notes, diagnosis and billing codes, orders, EHR chart summarization, patient after-visit summaries, and clinical decision support. In every instantiation the construction is identical: the clinician's review-and-attestation event is the ground truth, every accepted unit is work the system completed, and every correction is residual effort returned to the clinician. The primary contribution of this paper is the benchmark itself: the metric, its degenerate cases, and a reporting protocol under which ER claims are auditable and cross-system comparable. Alongside it we report an initial headline measurement from the documentation instantiation: over one million signed encounters across a production window exceeding six months and thirteen medical specialties, Knowtex's proprietary fine-tuned clinical foundation models operating inside a closed feedback architecture achieve an aggregate ER of 97.99%, with per-specialty aggregates spanning 96.8-98.9%. This release reports the protocol's checklist partially, and states which companion statistics are withheld; the benchmark is offered so that this figure, and every figure reported after it, can be held to the same standard.