EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards

2026-08-24Artificial Intelligence

Artificial Intelligence
AI summary

The authors created EarthVerse, a benchmark to test how well scientific AI agents can analyze complex Earth system events like natural hazards. It includes 405 tasks based on real events where agents must gather and use different types of evidence, do calculations, and explain their work clearly. They tested 25 AI systems and found that although many can handle individual steps accurately, very few keep everything consistent throughout the whole analysis process. This shows current AI still struggles to fully understand and connect all parts of complex Earth data. EarthVerse aims to help measure and improve the reliability of these AI methods in Earth science.

Earth-system analysisNatural hazardsScientific agentsBenchmarkEvidence reconciliationProvenanceTask evaluationTool-using protocolAnswer-unit accuracyPhysical interpretation
Authors
Zhiqing Cui, Xinxiang Yin, Yihong Tang, Xinglang Zhang, Yuanzhe Hu, Siru Zhong, Weidong Tang, Yuxuan Liang, Weijia Li, Ming Jin, Shirui Pan, Yuhao Kang, Dingyi Zhuang, Jinhua Zhao
Abstract
Earth-system analysis reconstructs changing physical processes from observations that differ in source, scale, timing, and modality. Natural hazards make this work consequential because incomplete evidence can change estimates of severity, exposure, and mechanism. We introduce EarthVerse, a benchmark that evaluates scientific agents through package-scoped investigations. Its 405 reproducible tasks are grounded in 199 documented events and 19 hazard families. Agents inspect heterogeneous event packages, choose compatible evidence, execute transparent calculations, reconcile source differences, and preserve provenance in the final answer. We provide executable ground truth that decomposes each task into fine-grained answer units, together with task-specific rubrics that assess the supporting research process while allowing multiple valid paths. We evaluate 25 model and agent systems under a controlled tool-using protocol, then use controlled studies to locate failures in evidence access, tool selection, memory, reasoning, interaction, and scientific execution. Across systems, the best mean answer-unit accuracy is 84.65%, while the highest Strict@95 is only 34.81%. The gap shows that current agents often complete individual steps without maintaining a consistent chain across evidence, scales, units, calculations, and physical interpretation. EarthVerse provides a reproducible basis for measuring end-to-end scientific reliability in dynamic Earth systems.