Training-free medical knowledge approach improves disease diagnosis accuracy

Training-Free Clinical Reasoning through Medical Ontologies and Cognitive Mapping: A Symbolic-Probabilistic Knowledge Graph Framework

Artificial Intelligence

Summary

Medical diagnosis often relies on systems trained on data with known outcomes. This paper introduces a new way to reason about diagnoses without training on outcome labels, using existing medical knowledge mapped to patient data. The method separates and ranks evidence related to diseases, allowing classification of patient groups based on how well their symptoms match known disease patterns. The authors tested their approach on several diseases like dengue and malaria, showing high accuracy without needing labeled training data.

What this means in practice

  • For hospital data teams: Use a label-free knowledge-driven system to audit and interpret diagnostic evidence across patient cohorts for diseases like dengue and malaria.
  • For healthcare software developers: Implement diagnostic support tools that combine symbolic medical ontologies with patient data without needing supervised training datasets.

Authors

Surajit Das

Abstract

Most clinical prediction systems learn patient-variable-outcome associations; we investigate a training-free diagnostic paradigm mapping patient observations to explicit medical knowledge. CKG Reasoner integrates candidate-specific Evidence Feature Nodes, patient-reference matching, a bounded Information Gate, knowledge-weighted evidence accumulation, disease similarity, and decisive clinical rules. Missing-aware normalization and coverage auditing distinguish absent from unavailable evidence. Candidate ranking is separate from outcome-label-independent K-means clustering, which uses four derived evidence coordinates (evidence strength, relative magnitude, directional similarity, and evidence completeness), not raw predictors or targets, to derive cohort-level assignments. Across six retrospective cohorts - four dengue (N = 1000, 1523, 989, 1018), malaria (N = 2190), and influenza (N = 4569) - a uniform, label-free, cohort-fitted K = 2 protocol yielded positive-class F1 scores of 0.996, 0.634, 0.936, 0.917, 0.695, and 0.842, and all-record accuracies of 0.996, 0.558, 0.914, 0.893, 0.707, and 0.906, respectively, with full partition-decision coverage using the frozen package and disease-specific knowledge representations. Neither scoring nor clustering uses outcome labels. Logistic regression provides a supervised baseline. Influenza incorporates confirmatory molecular PCR and is not independent pre-test prediction. Results characterize knowledge-grounded evidence separation, auditability, and sensitivity, not prospective clinical validity or comparative superiority. FOL/LLM-based clinical explanation remains unevaluated.