Science sandboxes measure the scientific capability of AI agents

2026-08-31Artificial Intelligence

Artificial Intelligence
AI summary

The authors introduce 'science sandboxes,' a way to test how well AI can learn scientific rules by experimenting, getting feedback, and updating ideas. These sandboxes let AI try different kinds of experiments, from real physical tests to using computer models and made-up rules. They tested this approach in biology areas like gene regulation and protein fitness. The authors found that AI could improve numbers without truly understanding the rules, especially when dealing with unfamiliar biological systems. This framework helps measure scientific thinking in AI and shows where it struggles.

AI agentsscientific reasoningexperimentationregulatory genomicsprotein fitness predictionempirical datahypothesis revisionmachine learningbiological priors
Authors
Arya S. Rao, Rodrigo I. Castro, Sager J. Gosai, Kenneth B. Hsu, Yasha Ektefaie, Shantanu Singh, Sangeeta N. Bhatia, Steven K. Reilly, Ryan Tewhey, Eric S. Lander, Pardis C. Sabeti
Abstract
Scientific progress depends not only on finding solutions, but on learning the rules that explain why they work and using that understanding to design better experiments. We introduce science sandboxes, a framework for studying this capability in AI agents through repeated cycles of experimentation, feedback, and hypothesis revision. Science sandboxes invite an agent to query the natural world in different ways, ranging from "wet" physical experiments, to "damp" predictive models trained on empirical data, to "dry" invented rules. By establishing a common experimental loop and a protocol for evaluating agents within it, science sandboxes allow assessment of both quantitative performance on specific metrics and qualitative scientific reasoning, across a spectrum of empirical verifiability. Here, we instantiate this framework in two biological settings, models of regulatory genomics and protein fitness prediction, and examine the capabilities of frontier agents. Across these settings, we could see when agents successfully optimized a quantitative metric without understanding the rules underlying the system. In particular, their scientific reasoning deteriorated when they encountered systems whose rules fell outside familiar biological priors. By highlighting such failure modes, science sandboxes make the frontier of scientific capability measurable and provide a controlled setting in which to study and ultimately expand it.