Machine learning tool finds hidden errors in atomic models beyond tests
MLIP Detective: Active Failure Mode Discovery Beyond Benchmark Scores for Machine-Learning Interatomic Potentials
Machine Learning
Summary
Machine-learning models that predict how atoms interact are tested using standard benchmarks, but these tests can miss some critical mistakes. The authors created a tool called MLIP Detective that uses physics knowledge to actively search for these hidden errors. It examines unusual cases through quick simulations and only alerts experts when serious problems appear. Using this tool, they discovered that a popular model incorrectly predicted certain chemical systems to be less stable than they should be, likely due to issues in its training data.
machine-learning interatomic potentialsbenchmark evaluationfailure mode discoveryphysics-informed searchadsorbate-surface systemsenergy predictionsimulation screeningtraining datamodel validationactive testing
Authors
Ryuhei Okuno, Nontawat Charoenphakdee, Kaoru Hisama, Yuta Tsuboi
Abstract
Universal machine-learning interatomic potentials (u-MLIPs) aim to generalize across diverse configurations. Benchmarks enable reproducible evaluation but may not expose failures outside their predefined scope. Here, we show that physics-informed search can complement benchmark-based evaluation by uncovering hidden failure modes. We introduce MLIP Detective, an agentic framework for active failure mode discovery. Starting from benchmark evidence, MLIP Detective generates falsifiable, physics-informed failure hypotheses, screens them with inexpensive simulations, and escalates only the most suspicious cases to human experts together with proposed verification protocols. Without issue-specific prompting, MLIP Detective identified and characterized a systematic anomaly in MACE-MPA-0: the model predicted some relaxed adsorbate-surface systems involving O- or F-containing adsorbates to be higher in energy than their corresponding separated fragments. Using cross-model comparisons, MLIP Detective further inferred a likely training-data origin for the anomaly, consistent with recent reports.