onepot-Bench 0: towards lab-aware in silico chemistry benchmarks
2026-08-03 • Machine Learning
Machine Learning
AI summaryⓘ
The authors created onepot-Bench 0, a new test to see how well language models understand and help with chemistry tasks in real labs. Their test has three parts: one checks basic chemistry knowledge and math skills without extra tools, another tests how models handle safety and refusal for different chemicals, and the last one looks at predicting chemical reactions and choosing catalysts using real lab data. This helps measure if models can reliably assist in physical experiments, beyond just knowing facts. The authors' benchmark focuses on skills needed to work safely and effectively in a real chemistry lab.
language modelscheminformaticssynthetic chemistryreaction predictioncatalyst selectionbenchmarkssafety evaluationnumerical reasoningwet-labexperiment planning
Authors
Brandon Wang, Andrei S. Tyrin, Daniil A. Boiko
Abstract
Language models are playing an increasingly important role in laboratory science, performing tasks such as experiment planning, execution, and post-hoc analysis. However, precisely measuring their abilities is difficult, as scientific capabilities require a mixture of both problem-solving skills and domain-specific intuition. Existing evaluations rarely measure the capabilities required to make reliable decisions in a physical laboratory and often rely on public data that may have appeared in model training corpora. We introduce onepot-Bench 0, a proprietary benchmark suite for evaluating language models on synthetic chemistry capabilities relevant to wet-lab execution. onepot-Bench 0 comprises three complementary evaluations: ChemAbacus measures tool-free cheminformatics literacy and numerical reasoning; SynthRefusal characterizes safety and refusal behavior across a variety of benign, controlled, and designer-drug targets; and SynthBench evaluates reaction-outcome prediction and catalyst selection using private experimental data generated in our laboratory. Together, these evaluations probe basic competency, reliability, and deeper knowledge, all skills which are required for reliable performance in the lab.