MolSC improves molecular models by tracking small chemical changes

MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs

Machine LearningArtificial Intelligence

Summary

Molecular models using language AI have trouble understanding how tiny changes to molecules affect their behavior. The authors created MolSC, a large dataset that shows how adding small parts to molecules changes their properties. Training AI models on MolSC helps them predict these changes more accurately. This improves the models’ ability to understand detailed chemical effects and perform better on related tasks.

What this means in practice

  • For drug discovery teams: Predict how small chemical modifications alter drug molecules’ bioactivity and safety to guide molecule design decisions.
  • For chemical informatics engineers: Enhance molecular AI tools with fine-grained predictions of property changes driven by substituents for better molecular analysis.

Authors

Hyuntae Park, Sooyeon Kim, Jiwon Park, SangKeun Lee

Abstract

Recent advances in natural language processing have led to molecular Large Language Models (LLMs) with strong performance across diverse chemistry tasks. However, they still struggle to capture fine-grained structure-property relationships, particularly how small, localized modifications alter a molecule's behavior. To address this limitation, we introduce MolSC, a dataset of substituent contributions, defined as property changes induced by attaching specific substituents to molecular scaffolds. Curated from manually annotated bioactivity records, MolSC spans structural-alert liability, target-specific bioactivity, and physicochemical descriptors, and contains 181K substituent-level examples for training. We further propose MolSC-Bench, a held-out evaluation benchmark of 1,541 examples disjoint from MolSC at the scaffold, substituent, and molecule levels. Our experiments show that existing molecular LLMs and strong proprietary models such as GPT-5.2 and Gemini-3-Flash show limited reliability in substituent contribution prediction. In contrast, training on MolSC substantially improves this ability and achieves strong performance across diverse downstream molecular tasks. These results highlight substituent contribution learning as a key component of fine-grained molecular understanding.