Benchmark reveals challenges in testing causal discovery models fairly
CausalArena: Benchmarking Causal Discovery in the Foundation Model Era
Machine Learning
Summary
Causal discovery tries to find cause-and-effect relationships from data, which is important for science and making decisions. The authors show that testing how well these methods work is tricky because different tests use different types of data and rules. They created CausalArena, a new testing system that includes a variety of data types to better evaluate these methods. Their experiments found that some models that do well in one test may not do well in others, highlighting the need for better testing methods. This helps us understand that evaluating causal discovery tools is more complex with today's advanced AI approaches.
causal discoverystructural causal modelsbenchmarkfoundation modelspretrainingsynthetic dataevaluation protocolcausal graphinterventionreal-world datasets
Authors
Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
Abstract
Causal discovery aims to uncover causal structures from data and is fundamental to scientific reasoning and intervention-based decision making. Its evaluation relies heavily on structural causal models (SCMs), which specify a causal graph together with the mechanisms that generate data, yet existing studies differ substantially in graph families, mechanisms, and evaluation protocols. The emergence of causal discovery foundation models (CDFMs) further complicates evaluation: performance may reflect not only causal discovery ability, but also overlap between pretraining environments and test SCMs, making results on fixed synthetic benchmarks difficult to interpret. We introduce CausalArena, a unified and evolvable benchmark for causal discovery under a common protocol. Synthetic SCMs supply controlled breadth over structures and mechanisms; semantic operational SCMs provide human-auditable, semantically grounded environments beyond standard synthetic generators; and formula-grounded SCMs test discovery under explicit scientific mechanisms. Public real-world datasets provide an additional external-validity check. Experiments across classical, neural, and pretrained methods reveal substantial ranking shifts across SCM families and protocols, showing that strong performance in one benchmark regime does not reliably transfer to others. These results highlight benchmark diversity and pretraining--evaluation overlap as central challenges for evaluating causal discovery in the foundation model era.