AlphaDiverse improves stock factor research with diverse local AI agents

AlphaDiverse: Post-Training Local Quantitative Research Agents for Diverse Exploration in Alpha Factor Mining

Artificial IntelligenceComputational Engineering, Finance, and ScienceMultiagent Systems

Summary

Finding new strategies to predict stock prices often relies on repeated ideas, limiting success. The authors created AlphaDiverse, a system where multiple AI agents generate varied research paths and learn from them locally, rather than relying on external tools. This approach encourages exploring a wider range of possibilities and reduces repeated focus on a few popular strategies. Their experiments with Chinese stocks show that AlphaDiverse balances good predictions with more diverse research outcomes.

What this means in practice

  • For quantitative finance teams: Develop local AI agents that explore diverse stock prediction strategies beyond common paths using AlphaDiverse’s multi-agent system and post-training.
  • For machine learning engineers: Build AI research systems that improve creativity by fine-tuning agents on diverse research paths and optimizing them jointly with feedback-driven methods.

Authors

Qingzhuo Wang, Zikun Wei, Zhihua Wei, Wen Shen

Abstract

Large language model (LLM)-based multi-agent systems can automate alpha factor mining, but their reliance on external APIs limits control over cost, availability, and confidentiality. Long research loops also tend to revisit a few successful economic mechanisms that lead to research path collapse. To address these limitations, we propose AlphaDiverse, a framework that integrates a multi-agent alpha research system, diverse research path collection, and post-training for local agents. We let the research system generate complementary plan portfolios and vary research environments across loops to collect diverse research paths. Using these diverse traces, we warm-start local Planner and Realizer agents with supervised fine-tuning. Then, we propose a joint GRPO method to optimize both of them using predictive quality and diversity of contributions. Research feedback is confined to inner period data, while a frozen final model is evaluated on a later outer period data, thereby avoiding test-set tuning. Experiments across four Chinese stock universes show that AlphaDiverse can combine competitive prediction with broader exploration.