ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions
2026-08-31 • Artificial Intelligence
Artificial IntelligenceComputation and Language
AI summaryⓘ
The authors created ScienceArena, a new test made from real science competitions in physics, chemistry, and biology, to better check how well large language models (LLMs) can solve tough multi-step science problems. They carefully digitized and verified official exam materials with expert help and used expert scores to train LLMs to grade answers reliably. When testing 14 recent LLMs, some performed well enough to earn medals on certain exams, but chemistry questions and complex multi-step problems were still hard. The authors also made an online demo for others to try.
large language modelsolympiad-style benchmarkscience competitionsphysicschemistrybiologyprocess-credit rubricvisual groundingLLM evaluationinteractive demo
Authors
Guangxiang Zhao, Qilong Shi, Xusen Xiao, Wenpu Liu, Yaoming Li, Linfeng Hao, Shuyang Hou, Zijian Guo, Xinrui Zhang, Yuntian Zhao, Zhengyang Wang, Wenrui Liu, Yuhan Wu, Tong Yang, Lin Sun, Xiangzheng Zhang
Abstract
Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in physics, chemistry, and biology, including IPhO and IChO 2025--2026, IBO 2023, USAPhO 2026, and USNCO 2025. Its open-ended, multi-step problems use process-credit rubrics, making faithful scoring difficult. We build ScienceArena through an expert-audited digitization pipeline that converts official exams, figures, solutions, and rubrics into structured items verified by olympiad medalists. To scale evaluation beyond costly human grading, we calibrate LLM-as-judge against medalist ground truth on archived answers from five models across IPhO and IChO; two strong judges stay within one point of expert total scores. Medalist notes show that failures often stem from visual grounding, structure fidelity, and global problem control rather than missing terminology. Evaluating fourteen recent LLMs with interleaved solving, we find that top models obtain medal-equivalent rubric scores on several public international exams, while chemistry and long-horizon consistency remain key bottlenecks. We provide an interactive \href{https://science-arena.onrender.com/}{demo}.