SymbolicArena streamlines evaluation for symbolic regression benchmarks
SymbolicArena: A Unified Infrastructure for Benchmark Distillation and Dynamic Evaluation in Symbolic Regression
Machine Learning
Summary
Finding simple math formulas that explain data, called symbolic regression, is useful but hard to test because existing sets of problems are either too big to check thoroughly or too small to be reliable. The authors created SymbolicArena, a system that shrinks a large collection of tasks into a smaller, well-chosen set that still tests different methods fairly. This smaller set, Core50, lets researchers compare algorithms faster without losing important differences. They also found that current methods often fit data well numerically but struggle to uncover true underlying formulas.
What this means in practice
- •For machine learning engineers: Evaluate and compare symbolic regression algorithms efficiently using a smaller, validated benchmark without sacrificing test quality.
- •For software developers in scientific computing: Integrate a unified platform that runs different symbolic regression methods under a consistent protocol to simplify method testing and debugging.
Authors
Ziwen Zhang, Xiju Wu, Yuheng Jing, Runxiang Wang, Boxiao Wang, Yifan Zang, Yifan Zhang, Yang Wang, Kai Li, Yifan Zhang, Huilin Xu, Jian Cheng
Abstract
Symbolic regression (SR) seeks concise and interpretable mathematical expressions from data for scientific equation discovery. Existing SR benchmarks face a tradeoff between evaluation cost and benchmark validity. Repeated evaluation of large task pools is expensive, and compact benchmarks lack systematic evidence of preserved task diversity and algorithm discriminability. SymbolicArena provides a unified infrastructure for benchmark distillation and dynamic evaluation. The framework standardizes 664 heterogeneous tasks with executable ground truth expressions and distills the Full Task Set into Core50, a validated benchmark of 50 tasks. The distillation process preserves task coverage and algorithm discrimination under explicit balance constraints. SymbolicArena applies a unified execution protocol to heterogeneous SR algorithms and produces comparable outputs and search trajectories. Multi Axis Evaluation characterizes numerical quality, symbolic quality, and search behavior. Core50 reduces evaluation workload by 92.5% and maintains agreement with Full Task Set evaluations. Experiments show that SymbolicArena achieves 72.6% to 86.7% lower approximation error than alternative selectors, further supporting its fidelity to the Full Task Set. Evaluation reveals a substantial gap between numerical fitting and symbolic recovery across current SR methods, suggesting that reliable equation recovery remains an open challenge.