Higher-order grammar improves molecule generation and learning quality

Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry

Machine LearningArtificial Intelligence

Summary

Molecules have complex shapes that are hard for computers to understand fully, especially parts like rings or common patterns. The authors created a new way called Higher-order Grammar Representation that breaks molecules down into simple rules capturing these complex parts. This method works well with existing sequence-based computer models, making molecule generation and learning more accurate and efficient. They also made a new big ring-focused set of molecules to better test such methods. Their approach leads to perfect molecule validity in generation and better results in predicting molecular properties.

What this means in practice

  • For drug discovery teams: Generate valid and diverse molecular structures with complex ring systems for drug candidate design using higher-order grammars.$Commercial implications: Enables novel molecule creation platforms targeting pharmaceutical companies with improved chemical validity and diversity.
  • For chemical informatics engineers: Train molecular property prediction models with enhanced representations that capture complex molecular topology for better prediction accuracy.

Authors

Yiming Huang, Yujie Zeng, Vijay Prakash Dwivedi, Simone Foti, Jianmin Wang, Jure Leskovec, Tolga Birdal

Abstract

Molecular learning models are strongly shaped by their underlying representations. Yet standard sequential and graph formalisms struggle to explicitly encode higher-order topology, such as ring systems and recurring motifs. Existing higher-order representations can capture these structures directly, but they are often computationally demanding and difficult to decode into valid molecules. Here, we introduce Higher-order Grammar Representation (HGR), a principled, topology-aware framework that lifts molecules to combinatorial complexes and parses each complex into a compact sequence of production rules under a context-free higher-order grammar. By serialising higher-order topology into rule sequences, HGR makes these structures directly compatible with standard sequence models, avoiding the computational overhead of explicit higher-order encodings while preserving topological expressiveness. To reduce benchmark bias towards simple ring systems, we construct RingDiv, a ring-enriched benchmark containing 1.18 million molecules, including the curated RingDiv300k subset, and introduce the ring diversity index (RDI) to quantify ring-system coverage. In molecular generation, HGR-based models uniquely combine 100% validity by construction with leading distributional alignment, ranking first in FCD on all five generation benchmarks. In representation learning, HGR-FM achieves the highest mean AUC across seven MoleculeNet benchmarks under both transfer protocols, improving on the strongest baseline by 8.3 and 3.3 AUC points under probing and full fine-tuning, respectively. Collectively, these results establish HGR as an efficient higher-order representation for molecular generation and transferable representation learning.