What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation
2026-08-24 • Computation and Language
Computation and LanguageArtificial Intelligence
AI summaryⓘ
The authors created a new test called Lit2Test to better judge research ideas proposed by language models. Instead of vague opinions or later paper comparisons, Lit2Test requires each idea to state what specific result would prove it wrong, making evaluation clearer. They tested four advanced models using many comparisons judged fairly and found clear differences based on how well the models suggested tests, not just how nicely they wrote. The authors also made their test and methods publicly available for others to use.
large language modelsbenchmarkresearch proposalsfalsifiabilityevaluation protocolpairwise comparisonsmodel rankingreliability audit
Authors
Ziyue Wang, Aomufei Yuan, Yiran Yao, Linli Yao, Hongyao Zuo, Ziwen Gong, Yuanxin Liu, Shicheng Li, Yishuo Cai, Tong Yang, Xu Sun, Xiaohui Li, Haoli Bai
Abstract
Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decision rule: free-form judging sways with style and position, and scoring against a later paper rewards recovery of one realized trajectory. We introduce a benchmark that carries a proposal from Literature to Test: the Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable. Built prospectively from 200 real-paper neighborhoods, Lit2Test elicits proposals from four frontier models and compares them through 1,200 pairwise comparisons judged blind in both presentation orders. The protocol audits its own reliability through diagnostic controls and bounded human calibration, with three annotators corroborating the conclusions within explicitly stated reliability bounds. Lit2Test recovers a strict ranking of the four models in all 10,000 bootstrap replicates, and the separation comes from the quality of the proposed tests and metrics rather than from surface fluency. We release the benchmark, construction pipeline, and audit artifacts for public use.