Papers for

clinical operations teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Clinical AI system efficiency measured by effort reduction across specialties

KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI

Abstract: Clinical AI systems are evaluated with instruments built for research settings (reference-based similarity metrics and expert rubric panels) that measure resemblance to an artifact rather than reduction of a burden. We introduce KnowBench, pioneered by Knowtex, whose unifying metric is Effort Reduction (ER): the proportion of system-generated clinical work product accepted by the responsible clinician under expert and safety review. ER is defined once and instantiated per task across the administrative workload clinical AI automates: visit notes, diagnosis and billing codes, orders, EHR chart summarization, patient after-visit summaries, and clinical decision support. In every instantiation the construction is identical: the clinician's review-and-attestation event is the ground truth, every accepted unit is work the system completed, and every correction is residual effort returned to the clinician. The primary contribution of this paper is the benchmark itself: the metric, its degenerate cases, and a reporting protocol under which ER claims are auditable and cross-system comparable. Alongside it we report an initial headline measurement from the documentation instantiation: over one million signed encounters across a production window exceeding six months and thirteen medical specialties, Knowtex's proprietary fine-tuned clinical foundation models operating inside a closed feedback architecture achieve an aggregate ER of 97.99%, with per-specialty aggregates spanning 96.8-98.9%. This release reports the protocol's checklist partially, and states which companion statistics are withheld; the benchmark is offered so that this figure, and every figure reported after it, can be held to the same standard.

Mon 14 SeptArtificial Intelligence
The gist
Current ways to measure how well clinical AI systems work focus on how similar their output is to examples, not on how much they actually help doctors save time. The authors created KnowBench, a benchmark that measures how much a clinical AI reduces the work doctors have to do by checking how much of the AI’s work is accepted without changes. They tested it across many clinical tasks and found their AI models reduced doctors’ effort by about 98% overall. This new way to measure AI’s usefulness helps compare different systems fairly based on the effort saved.
Open 2609.15794v1