Agentic symbolic regression improved by structured runtime support
SRHarness: A Harness for Agentic Symbolic Regression
Artificial IntelligenceMachine LearningSymbolic Computation
Summary
Symbolic regression is a way for computers to find mathematical formulas that explain data. The authors introduce SRHarness, a tool that helps manage the steps and information involved in this process, making it more efficient and accurate. They show that by organizing scientific actions and keeping track of hypotheses better, the system finds correct formulas more often than previous methods. This improvement holds true even when the descriptions or variable meanings are hidden, proving its strength.
What this means in practice
- •For machine learning engineers: Deploy agentic symbolic regression tools with enhanced runtime support to improve formula discovery accuracy over complex datasets.
- •For data analytics teams: Use structured scientific state management to maintain and explore hypothesis trajectories, boosting interpretability in model-building workflows.
Authors
Zihan Yu, Shixuan Zhou, Hao Huang, Jingtao Ding, Yong Li
Abstract
Recent agentic symbolic regression approaches increasingly rely on large language models to analyze data, select scientific operations, and refine hypotheses over long search trajectories. In such systems, performance depends not only on the underlying model and search strategy, but also on the runtime infrastructure that supports scientific search. We introduce SRHarness, a domain-specific harness for agentic symbolic regression built around three mechanisms: composable scientific actions that provide a common interface over raw, transformed, and candidate-derived quantities; persistent scientific state that retains evaluated hypotheses and exposes compact model-facing views; and trajectory lifecycle management that coordinates continuation, branching, restart, and termination. On LLM-SRBench, SRHarness consistently improves both numerical generalization and symbolic recovery under matched LLM backbones. With DeepSeek-v4-flash-0731, it achieves 93.69% symbolic accuracy on LSR-Transform, compared with 62.16% for SR-Scientist, and retains 72.97% accuracy on an anonymized variant that removes scientific descriptions and variable semantics, versus 39.64% for SR-Scientist. Under the same DeepSeek-v4-flash-0731 backbone, SRHarness also substantially outperforms Codex (72.97% vs. 20.72%) and reaches performance comparable to Codex with GPT-5.5, while simply providing Codex with the same scientific tools does not reproduce this advantage. These results show that effective agentic symbolic regression depends not only on models or tools, but also on structured runtime support for organizing scientific actions, accumulated hypotheses, and long-horizon search.