Scientific idea generation improves with active ai exploration
AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era
Artificial IntelligenceComputation and Language
Summary
Coming up with new scientific ideas is important for making discoveries, and AI systems can help with this. The authors created AgentIdeaBench, a test to see how well AI models generate ideas both by just reading papers and by actively searching for information. They found that AI models do much better when allowed to explore actively, showing bigger improvements in creating clear and feasible ideas. Stronger AI models benefit more from this interactive approach, while some methods help medium-level models refine ideas. This benchmark helps measure how AI can support scientific creativity in a more realistic way.
scientific ideationlarge language modelsactive explorationbenchmarkhypothesis generationretrievalreasoningagentfeasibilityscientific world modeling
Authors
Yunxiang Mo, Tianshi Zheng, Yisen Gao, Rui Wang, Newt Nguyen Kim Hue Nam, Kelvin Kiu Wai Tam, Jiaxin Bai, Yangqiu Song, Ginny Wong, Simon See
Abstract
Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and autonomous AI scientists depend on it. Existing evaluations largely assess it by asking models to generate ideas from a static, curated set of reference papers. That passive setup departs from the retrieval-and-reasoning workflow of modern AI scientists, and it becomes less discriminative as models improve. We introduce AgentIdeaBench, a multidisciplinary benchmark that evaluates scientific ideation under two matched settings, static observation and active exploration. We report matched Static-Active evaluations for 33 LLMs across 40 densely scored subfields spanning five disciplines, using a multidimensional, literature-verified scoring framework whose critics assess originality against retrieved prior art. Active exploration reveals considerably more capability headroom, and that headroom is unevenly distributed across models. Performance scales about twice as fast as under static observation, and the exploration gain is capability-gated, favoring the strongest models over the weakest. The gain reflects better grounding, improving feasibility, clarity, and specificity while leaving measured originality unchanged under our critics. We further explore Scientific World Modeling, a generation-time loop that refines a draft hypothesis through structured thought experiments. It benefits mid-capability models, and its impact diminishes among frontier models that appear to have internalized such reasoning patterns already. AgentIdeaBench gives future work on scientific ideation a measurement basis suited to the agent era.