Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

2026-07-20Machine Learning

Machine LearningArtificial IntelligenceComputation and Language
AI summary

The authors study how well large language models (LLMs) can design drug molecules by understanding 3D shapes of proteins and other spatial rules, compared to specialized diffusion-based models. They introduce a new test called 3D-Fit that checks if LLMs can create molecules fitting into protein pockets while following multiple 3D constraints at once. Their results show that although LLMs are not yet as good as specialized methods, they show promise in handling several spatial conditions simultaneously. This suggests LLMs might improve and be useful for complex molecular design in the future.

structure-based drug design3D molecule generationlarge language modelsdiffusion modelsprotein pocketpharmacophoreligandspatial constraintsmolecular dockingbenchmarking
Authors
Thomas MacDougall, Maksim Kuznetsov, Roman Schutski, Rim Shayakhmetov, Maxim Malkov, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
Abstract
Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradigm for high-quality 3D molecule generation, LLM-based methods are rapidly emerging in molecular design and have shown competitive performance in pocket-conditioned molecular generation. However, their ability to reason about physics and 3D spatial environments is largely underexplored. In this work, we systematically analyze whether current general-purpose LLMs are capable of navigating complex 3D constraints compared to established baselines such as specialized diffusion models. We consider 3D ligand generation conditioned on protein pockets together with ligand- and interaction-derived spatial constraints, including anchor fragments, pharmacophore points, and mandatory pocket-ligand interactions. To enable this evaluation, we introduce 3D-Fit - a token-efficient benchmarking strategy for assessing LLM performance on multi-conditioned spatial molecule generation. Our findings reveal a clear pattern in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.