Large language models improve traditional chinese medicine prescriptions with safer reasoning
Syndrome, Synergy, and Safety: Structured Reasoning and Knowledge-Driven Alignment for TCM Prescription Generation
Computation and LanguageArtificial Intelligence
Summary
Large language models can create Traditional Chinese Medicine (TCM) prescriptions, but they often miss important reasoning steps and safety rules. The authors found three main problems: no clear reasoning process, ignoring changes in patient follow-ups, and failing to avoid harmful ingredient combinations. They built a multi-stage method that teaches models to explain their reasoning, consider patient history, and follow strict safety rules. This approach made the models produce better and safer TCM prescriptions than previous versions.
What this means in practice
- •For medical ai developers: Develop AI tools that generate TCM prescriptions with clear diagnostic reasoning and follow patient treatment changes over time.
- •For healthcare software engineers: Build safer clinical decision support systems that enforce absolute contraindication rules in TCM prescription generation.
Authors
Zheng Chen, ZhiCheng Du, Haoxuan Li, Peiwu Qin
Abstract
Applying large language models to Traditional Chinese Medicine (TCM) prescription generation reveals three clinically critical gaps: models produce end-to-end mappings without auditable reasoning following the li-fa-fang-yao paradigm (SR Gap), treat each encounter in isolation without follow-up adjustment via sui zheng jia jian (LA Gap), and fail to enforce absolute contraindication rules such as Shi Ba Fan (SC Gap). We propose a progressive four-stage framework (SFT $\to$ PG-CoT $\to$ Dynamic $\to$ K-RL) that addresses each gap: PG-CoT constrains CoT distillation under the li-fa-fang-yao paradigm to produce auditable diagnostic chains, Dynamic SFT models patient trajectories with explicit transition reasoning, and K-RL encodes deterministic pharmacological rules as rule-based DPO preference signals. Across 12 fine-tuned models and 6 zero-shot baselines, our framework substantially improves prescription quality over zero-shot baselines---with a 7B model (Mistral-7B) surpassing zero-shot GPT-5 on all three TCM evaluation metrics.