Large language models improve reasoning by training on evolving problems

Frontier Learning: Training LLM Reasoners at the Edge of Capability

Machine LearningArtificial IntelligenceComputation and Language

Summary

Large language models (LLMs) can get better at reasoning by learning from practice problems. Previous methods trained these models on a fixed set of problems, but the models quickly outgrow them. The authors propose a new way called frontier learning, which keeps generating new, slightly harder problems as the model improves. This approach helps the model learn more effectively by always challenging it at the edge of its current ability.

What this means in practice

Authors

Robin Faro, Shyam Sundhar Ramesh, Ilija Bogunovic, Aurelien Lucchi

Abstract

Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO loss. This is fundamentally limiting, as learning signal arises only when policy rollouts mix successes and failures, causing the useful portion of any fixed pool to quickly become stale as the model improves. To address this, we propose frontier learning, an open-ended post-training approach in which procedural generators are used online to continually produce informative training problems. It treats the generator's task-specific parameters as a search space and uses a regret signal to prioritize and explore frontier difficulty levels in order to focus training at the edge of the model's evolving reasoning capabilities. Across several reasoning tasks and model families, our approach consistently achieves higher relative gains over fixed-pool baselines, demonstrating that effective post-training requires not only selecting useful problems, but continually generating them at the edge of capability.