Nereus improves large language model training efficiency on GPU clusters

Nereus: Adaptive Parallelism for LLM Post-Training

Distributed, Parallel, and Cluster ComputingArtificial Intelligence

Summary

Large language models (LLMs) often need a process called reinforcement learning post-training to improve their behavior. This process runs across many GPUs but can slow down when resources or workloads change during training. The authors designed Nereus, a system that adapts how the training work is spread across GPUs in real-time to keep things running efficiently. By doing this, Nereus reduces the waiting time for each training step and increases overall speed without much overhead. This helps train large models faster and more cost-effectively.

What this means in practice

  • For machine learning engineers: Run reinforcement learning post-training of large language models more efficiently by dynamically adjusting GPU usage during training.
  • For cloud infrastructure operators: Improve utilization and throughput of GPU clusters hosting large-scale language model training jobs by adapting resource allocation as workloads change.

Authors

Songlin Jiang, Tuo Shi, Sitong Zhang, Zeke Wang, Mario Di Francesco, Bo Zhao

Abstract

Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a consequence, an execution plan that was initially suitable can then become slow or even infeasible over time. However, adapting a job whose models share GPUs entails significant challenges: deciding whether a new plan is worth the transition cost, reusing the job's distributed state, and coordinating GPU transfers across models and stages. Nereus targets these challenges as a cost-aware runtime that adapts RL post-training jobs into efficient execution plans. Its low-overhead controller selects a memory-feasible global plan and admits the transition using a cost model calibrated against the running job. To estimate and execute a transition, Nereus represents the distributed state of each replica of a model-stage (one model in one stage) as an Elastic Model Unit. It then employs a global transition graph to order the transformations and GPU transfers of these units. In a trace built from real data, online TP/PP adaptation reduces average step latency by 27.7% relative to the initial fixed TP/PP layout with DP scaling. In a 1,000-step run reaching 1,024 GPUs, six transitions consume 0.079% of total run time. Nereus improves end-to-end 8B PPO throughput by 2.14--7.27$\times$ over OpenRLHF and by 1.10--1.47$\times$ over Verl across diverse clusters.