Quantitative Evidence Mining for Plausibility-Aware Biomedical AI
2026-08-31 • Computation and Language
Computation and Language
AI summaryⓘ
The authors explain that simply pulling out scientific claims from biomedical texts is not enough to make those claims trustworthy or useful. They argue that claims need to include detailed numbers, conditions, uncertainties, and sources to be properly understood and verified. Current AI methods often miss these important details, leading to claims that seem reliable but are hard to check or compare. The authors propose a new approach that treats extracted information as evidence with clear context and reliability checks, making AI outputs more transparent and trustworthy.
biomedical AIlarge language modelsknowledge graphsevidence synthesisquantitative evidence miningprovenanceuncertaintyplausibilityrelation extractionclinical trials
Authors
Negin Sadat Babaiha, Stefan Geissler, Marie-Christine Simon, Martin Hofmann-Apitius, Marc Jacobs
Abstract
Biomedical artificial intelligence (AI) systems increasingly extract, organize, and reuse scientific claims from literature, clinical trials, and regulatory documents. But automatic extraction alone does not make a claim reliable evidence: a claim becomes useful only when it can be traced to its source, linked to the quantitative details that support it, and read within its biomedical context and uncertainty. This matters as large language models (LLMs) and increasingly autonomous systems drive evidence synthesis, knowledge graph (KG) construction, and decision support. Many text-mining and LLM pipelines remain relation-centric: they capture entities and relations such as Drug--TREATS--Disease, but drop the dose, effect size, population, comparator, uncertainty, and conditions under which a claim holds. Such relations can look actionable yet remain hard to verify, compare, or reuse. In this perspective, we argue for a shift toward quantitative evidence mining---extracting values, units, measured entities and properties, context, uncertainty, provenance, and plausibility as structured evidence units that populate evidence-aware KGs and can be checked for source grounding, unit consistency, completeness, and biological plausibility. We outline a framework for plausibility-aware AI that treats extracted claims not as final answers but as auditable evidence objects, making clear what was measured, how much it changed, in which setting, with what uncertainty, and from which source. The central risk is not only incorrect extraction, but claims that look like evidence while lacking the structure needed to trust them.