PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data
2026-08-17 • Machine Learning
Machine LearningArtificial Intelligence
AI summaryⓘ
The authors created a method called PertMind that teaches language models to understand biological experiments by using data from how cells react when their genes are changed. Instead of relying on expensive human explanations, PertMind learns by getting rewarded when it correctly predicts gene responses. This helps the model improve at recognizing biological patterns and can perform various related tasks without extra training. Their work shows that using experimental data as learning feedback can help models develop useful biological reasoning skills.
large language modelsreinforcement learningcellular perturbation atlasgene expressionbiological reasoningsupervised learningphenotypic screeningpathway analysistransfer learningexperimental atlases
Authors
Zhenchao Tang, Xiaogang Xu, Tianxu Lv, Jiahui Guan, Jiale Zhou, Haohuai He, Zhi Song, Hanbo Huang, Jiehui Huang, Jiafei Wu, Zhe Liu
Abstract
Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces. Here we show that cellular perturbation atlases can instead become reinforcement-learning environments, where measured gene responses provide computable rewards for biological reasoning. We introduce PertMind, which combines trusted-trajectory supervised initialization with gene-, pathway-, and format-level reinforcement signals. Trained only on forward perturbation-response prediction, PertMind improved response inference in unseen cellular contexts while retaining general language capabilities. It also transferred without task-specific post-training to reverse perturbation identification, double-perturbation reasoning, phenotypic-screen prioritization, and biological-process interpretation. PertMind further generated biological profiles that supported competitive gene, cell, and donor representations across multiscale downstream tasks. These results support the hypothesis that reinforcement on experimental endpoints can concentrate reusable biological strategies already accessible to pretrained models. More broadly, perturbation-derived reinforcement learning offers a scalable route for transforming expanding experimental atlases into training environments for general-purpose biological reasoning.