ScientistTwo AI autonomously solves complex science problems
ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI
Artificial Intelligence
Summary
Solving new science problems means exploring what we don't know yet. The researchers introduce ScientistTwo, a smart AI that tackles scientific questions all by itself. It tests ideas, runs experiments, and even reviews its own work without people helping. Their tests show ScientistTwo writes expert-quality papers and produces working code that often beats human results. This AI acts like a scientific explorer, pushing knowledge forward without needing human guidance.
What this means in practice
- •For machine learning teams: Generate full experimental research papers and code that improve upon current AI models without human intervention.
- •For software development teams: Automatically produce verified, executable codebases for complex algorithmic problems guided by autonomous AI research cycles.
Authors
Jaehyun Nam, Jinsung Yoon, Yanzhou Pan, Yubo Wang, Rui Meng, Parthasarathy Ranganathan, Tomas Pfister
Abstract
Scientific discovery is defined by the ability to identify the boundaries of existing knowledge and venture into unexplored territory. The ultimate vision for AI in science is problem-driven autonomous research: given a fundamental challenge by a human expert, the AI independently navigates the scientific landscape, uncovers theoretical and empirical bottlenecks, and systematically expands the frontier of knowledge. In this paper, we introduce ScientistTwo, a fully autonomous multi-agent framework designed to realize this vision. Specifically, ScientistTwo takes an initial problem as input, establishes state-of-the-art baselines, formulates novel hypotheses, and coordinates specialized agents to orchestrate an end-to-end discovery cycle without human intervention. Moreover, the framework rigorously conducts experiments using diverse datasets and metrics, refines methodologies through automated ablation studies, and validates research findings via a closed-loop simulated peer-review rebuttal engine. To evaluate ScientistTwo's capabilities against the highest standards of human scientific achievement, we benchmark it across papers accepted at top-tier conferences such as ICLR, ICML, and NeurIPS. As a result, ScientistTwo autonomously generates expert-level, publishable papers and fully verified, executable codebases. Its solutions consistently outperform human state-of-the-art models, and achieve higher average review ratings than human-authored papers under automated AI review agents. These results show that ScientistTwo is not merely an assistive tool but an autonomous scientific pioneer capable of pushing the frontiers of human discovery. Project website: https://scientist-two.github.io/