Large language model agents accelerate inverse design of metal-organic frameworks for gas separation
2026-07-12 • Artificial Intelligence
Artificial Intelligence
AI summaryⓘ
The authors created LEMO Agent, a system that uses large language models to design new metal-organic frameworks (MOFs) for separating gases. Their method combines language-based idea generation with checks to ensure chemical validity and good separation performance. They tested it on methane/nitrogen and carbon dioxide/nitrogen separations and found it produces diverse, high-performing candidates better than other methods. Some of these candidates were then tested in lab experiments, showing that this AI-driven approach can help discover new MOFs more efficiently than traditional methods.
Metal-Organic Frameworks (MOFs)Gas separationInverse designLarge Language Models (LLMs)Transformer modelsMOFid standardizationGrand Canonical Monte Carlo (GCMC)LinkersTopologyChemical validity
Authors
Zhaolin Hu, Hehe Fan, Wangyihan Guo, Meng Xu, Chenhao Rao, Qiwei Yang, Yi Yang
Abstract
Metal-organic frameworks (MOFs) offer a highly modular platform for adsorptive gas separation, yet their vast reticular design space makes inverse design difficult under simultaneous constraints of chemical validity, separation performance, and structural diversity. Here, we present LEMO Agent, a large-language-model agent framework for closed-loop inverse design of gas-separation MOFs in MOFid space. LEMO Agent couples language-based candidate generation with MOFid standardization, explicit validity checking, Transformer-based property prediction, structured design memory, and multi-island exploration. Through iterative generate--validate--evaluate--remember cycles, the agent uses feedback from both successful and failed candidates to guide chemically constrained search across linker, metal, and topology choices. We evaluate LEMO Agent on CH$_4$/N$_2$ and CO$_2$/N$_2$ separation tasks. Compared with representative generative, optimization, and agentic baselines, LEMO Agent enriches high-performing candidates, improves predicted separation performance, and maintains broad chemical and topological diversity. Selected candidates are further reconstructed, evaluated by GCMC simulations, and passed through an experimental down-selection workflow based on chemical feasibility and ligand purchasability, leading to initial wet-lab synthesis and SEM characterization. These results demonstrate that large language model agents can serve as interpretable and scalable design engines for accelerating MOF discovery beyond conventional fixed-library screening.