Language model improves multi-goal chemical reaction optimisation
Dynamic language model representations for multi-objective reaction optimisation
Machine Learning
Summary
Optimising chemical reactions is tricky because chemists want the best mix of yield, safety, and other goals. The authors showed that by using a language model—an AI that understands text descriptions—they can represent reactions better than traditional methods. This new approach helps find the best reaction conditions faster and works well on different types of chemical reactions. They tested it on real laboratory experiments and achieved high-performing chemical results efficiently.
What this means in practice
- •For chemical process engineers: Accelerate optimisation of complex chemical reactions involving diverse catalysts and additives by using language model-based reaction representations.
- •For pharmaceutical formulation teams: Facilitate efficient multi-objective tuning of asymmetric hydrogenation reactions for drug synthesis by integrating language-based modelling with experimental automation.
Authors
Joshua W. Sin, David Ming Segura, Bojana Ranković, Siu Lun Chau, Marius D. R. Lutz, Andrea Anelli, Ryan P. Burwood, Kurt Püntener, Maximilian J. Notheis, Raphael Bigler, Philippe Schwaller
Abstract
Optimising chemical reactions across multiple objectives, such as yield, selectivity, and safety, is central to chemical synthesis, and model-driven approaches depend critically on how reaction components are represented. Established featurisations are either chemically uninformative, as with one-hot encodings, or, as with molecular descriptors, do not readily extend across chemically distinct components. For structurally and functionally diverse components, it is therefore unclear what a shared representation should contain. Constructing such a representation is itself a challenging research undertaking that must be revisited for each new reaction system. Here we bypass this step by learning the reaction representation dynamically from text. Textual descriptions of reaction conditions are encoded by a fine-tuned language model trained jointly with Gaussian process surrogates, yielding task-adaptive representations within a multi-objective Bayesian optimisation loop. Across nickel- and palladium-catalysed cross-couplings in both sequential and parallel experimentation regimes, this approach reaches optimisation convergence in fewer experiments than descriptor libraries or one-hot encoding. Applied prospectively to a palladium-catalysed cyanation spanning mixed ligand denticity and heterogeneous additives, and to a three-objective asymmetric hydrogenation across chiral iridium and ruthenium catalyst families, two rounds of high-throughput experimentation (192 reactions, under 3% of each design space) delivered conditions translating directly to gram scale in 94% and 84% isolated yield, the latter at 99.6% enantiomeric excess.