Ai agents struggle to find hidden scientific mechanisms beyond observable laws
MechBench: Can AI Scientific Agents Discover Mechanisms Beyond Phenomenal Laws?
Artificial IntelligenceMachine LearningSymbolic Computation
Summary
Finding scientific laws that match what we see is important, but understanding the hidden reasons behind those laws is even harder. The authors created MechBench, a test to see if AI can discover these hidden mechanisms, not just the laws. They found that current AI models can often find the observable laws but fail to identify the true underlying mechanisms. This shows that discovering mechanisms is a different and more difficult challenge than just finding the rules that describe what happens.
What this means in practice
- •For ai developers: Improve AI models that assist in scientific discovery by focusing on mechanism identification beyond fitting observational data.
- •For educational software developers: Create tools that help students explore not only scientific laws but also underlying mechanisms, improving conceptual understanding.
Authors
Zihan Yu, Jiadong Zhang, Jialin Cheng, Jingtao Ding, Yong Li
Abstract
Scientific discovery requires not only recovering mathematical laws that describe observable behavior, but also identifying the mechanisms that generate them. Existing benchmarks for symbolic regression and scientific agents primarily evaluate phenomenal-law recovery, leaving mechanism discovery largely untested. We introduce MechBench, a benchmark that explicitly separates these two capabilities. Each task is defined by a mechanistic model, a structured set of scientifically meaningful relations whose joint consequences entail an observable phenomenal law, while agents receive only observational data and scientific context. We evaluate mechanism recovery through mechanism probes, which query internal scientific consequences that cannot be inferred from the phenomenal law alone. To reduce reliance on memorized textbook mechanisms, we construct unfamiliar variants through controlled, scientifically interpretable mutations of canonical mechanisms, and screen for mechanistic indistinguishability to exclude ambiguous instances admitting comparable competing mechanisms. Experiments across representative scientific agents reveal a substantial phenomenal--mechanism recovery gap: for Codex with GPT-5.6-sol, phenomenal-law accuracy reaches 35.00% on the Core-set while mechanism accuracy is only 13.75%, with mechanism recovery failing in 64.29% of cases where the phenomenal law is correctly recovered. The gap widens as mechanisms become increasingly mutated, and even providing the correct phenomenal law leaves mechanism recovery below 50%. These results reveal a substantial generalization gap in mechanistic reasoning and establish mechanism discovery as a distinct challenge beyond recovering observable scientific laws.