LLMs improve industrial optimization by adapting diverse strategies

LLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale Optimization

Machine LearningComputation and Language

Summary

Solving big, complex optimization problems in industry is hard, and current AI methods mostly work on small textbook examples. The authors found that different solution methods work better for different types of problems and scales. They trained large language models to switch between these methods adaptively by encouraging diverse problem-solving strategies. Their approach performs better than previous fine-tuned models on real-world, large-scale tasks.

What this means in practice

  • For operations research teams: Improve solution quality for large-scale industrial scheduling and resource allocation by using adaptive AI solvers that switch strategies based on problem structure.
  • For enterprise logistics planners: Handle diverse and complex optimization tasks efficiently with AI models trained to adaptively select different solving methods.

Authors

Shihao Zhang, Weiting Liu, Siyu Shao, Yitian Chen, Jianfeng Feng, Dongdong Ge, Yinyu Ye

Abstract

Scaling LLM-based optimization from textbook-scale instances to real-world, industrial tasks remains a critical open challenge. Existing approaches are predominantly evaluated on small, self-contained textual problems and often commit to a solver-integrated paradigm, limiting their ability to handle the scale and structural diversity of practical optimization workloads. In this work, we propose a practical framework for training open-source LLMs to tackle real-world, industrial-scale optimization. We first show empirically that solver-integrated reasoning, exact combinatorial algorithm, and heuristic search exhibit complementary strengths across different problem structures and scales. Motivated by this, we introduce Strategy-Diverse Reinforcement Learning (SDRL), which trains LLMs as adaptive optimization meta-solvers. SDRL leverages this complementarity through a correctness-gated hierarchical diversity reward that promotes robust exploration across varying strategies and within each strategy, effectively preventing premature strategy collapse. We further introduce a mixed-format training scheme that jointly supports both self-contained textual problems and file-grounded instances. Across comprehensive evaluations, our framework outperforms existing fine-tuned methods and frontier models including DeepSeek-V4-Pro and GPT-5.5, both on average across benchmarks and on industrial-scale optimization tasks.