Scaling search improves AI agents in financial factor discovery
From Search to Research: Exploring Search Scaling in Autonomous Quantitative Factor Mining
Artificial Intelligence
Summary
Finding better ways to research financial data automatically is tricky. The authors studied how increasing the amount of searching an AI agent does helps it perform better in financial research tasks. They found that stronger AI models start off better, but searching more can help smaller models catch up. Also, searching many ideas at once works better than going one by one. This means future improvements need both better AI and smarter ways to spend computing time.
What this means in practice
- •For quantitative finance teams: Improve automated financial factor discovery by balancing AI model strength with search strategies to enhance research quality.
- •For ai system engineers: Design adaptive AI agents that allocate search efforts dynamically for better problem-solving in complex research workflows.
Authors
Kangcheng Deng, Hui Cai, Jiacheng Lu, Chester Zhongshu Qian, Rui Sun, Beidi Luan, Jing Li, Daxin Jiang, Zuo Bai
Abstract
Inference scaling has been shown to improve large language model (LLM) performance, and this principle naturally extends to autonomous LLM agents through increased search budgets, which we refer to as *search scaling*. Although prior work has characterized the mechanisms, scaling behavior, and performance limits of LLM inference scaling, much less is known about these questions in autonomous research. Therefore, we investigate how search scaling affects research performance and what mechanisms drive these gains using 50 quantitative factor-mining tasks grounded in financial research reports. Each task requires an agent to carry out an end-to-end research loop, from interpreting a hypothesis and implementing it in code to evaluating and iteratively refining the resulting factor. Across nine models, we examine how model capability, search depth, and search organization shape factor quality by tracing performance across varying budgets, transferring intermediate research states between models, and comparing different search strategies. We find that (1) initial performance is more strongly associated with model capability, while deeper search can narrow cross-model gaps; (2) model grafting shows that the early research state materially shapes final performance; and (3) parallel search outperforms sequential search under the same iteration budget, consistent with benefits from broader coverage of the search space. Further trajectory analysis shows that higher-performing models more effectively diagnose failures, revise search directions, and preserve the intended economic hypothesis when selecting candidates. These findings suggest that future progress in autonomous research will require stronger models together with adaptive policies for deploying test-time computation throughout the research process.