Papers for

deep learning platform teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Solved learning rate scheduling improves large language model training

SOLAR: A State-Driven Online Learning Rate Scheduler for LLM Pretraining

Abstract: Learning-rate (LR) scheduling plays a central role in large language model (LLM) pretraining, yet current practice still relies heavily on hand-crafted heuristics such as Warmup-Cosine-Decay and Warmup-Stable-Decay. Because these schedules are fixed in advance, they cannot adapt to evolving optimization dynamics. Online learned scheduling within the Learning to Optimize (L2O) framework offers a dynamic alternative, but remains brittle at LLM scale due to noisy signals, delayed feedback, and the risk of catastrophic divergence. We propose SOLAR (State-driven Online Learning rAte scheduleR), a stabilized framework for reliable online LR adaptation. SOLAR uses a base schedule as a reference and learns bounded, state-dependent residual corrections for individual parameter groups. Each correction re-anchors to the base at every step, allowing the policy to adapt the LR without relearning the warmup-decay profile. A lightweight state representation and progress-aware reward guide online learning, while a Circuit-Breaker restores training after rare unsafe actions. Across autoregressive language-model pretraining, SOLAR improves final perplexity over tuned static schedules and automatic LR tuners for dense models from 60M to 1B, AdamW and Muon, and two MoE settings up to 3B. Matched 130M controls show that adding base anchoring and action bounds improves a global PPO controller from 27.09 to 23.74 final PPL, while group-wise control reaches 22.87 on the same two seeds. A residual policy trained on a 60M proxy can also be frozen and reused at larger dense scales without target PPO updates, remaining effective across a fourfold base-LR range. These results establish SOLAR as a practical learned LR controller for LLM pretraining.

Mon 28 SeptMachine Learning
The gist
Training large language models (LLMs) requires adjusting how quickly the model learns, but current methods use fixed schedules that may not fit changing training needs. The authors propose SOLAR, a method that dynamically adjusts the learning rate during training based on the model’s current state, allowing safer and more effective updates. SOLAR smartly corrects the learning rate at each step without losing the benefits of established schedules and can be reused across different model sizes. It improves the model’s accuracy and stability compared to traditional fixed schedules and previous automatic tuning methods.
Open → 2609.34681v1