Pearl boosts agentic reinforcement learning throughput by resource adaptation

PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning

Artificial Intelligence

Summary

Agentic reinforcement learning involves repeatedly simulating tasks, which can be slow and expensive. The authors introduce PEARL, a system that smartly adjusts how it uses GPU resources and breaks down work steps to keep machines busy and speed up learning. PEARL predicts how long tasks will take and decides when to switch strategies to avoid wasting computing power. Their tests show that this approach significantly speeds up the process compared to fixed setups and other methods.

What this means in practice

  • For machine learning engineers: Increase training speed and GPU efficiency in reinforcement learning workflows using adaptive resource allocation and execution modes.
  • For high performance computing teams: Optimize GPU cluster usage by dynamically adjusting task concurrency and resource sharing for workloads involving large language model rollouts.

Authors

Jiaan Zhu, Wei Gao, Youhui Bai, Zewen Jin, Ju Huang, Siran Yang, Jiamang Wang, Lin Qu, Cheng Li

Abstract

Multi-turn rollout dominates the cost of agentic reinforcement learning (RL). Asynchronous execution and elastic GPU resources can accelerate this stage, but adding rollout replicas yields diminishing returns while training GPUs remain idle between updates. We observe that effective resource use also depends on the prefill--decode (PD) configuration. Both the choice between colocation and disaggregation and the optimal PD ratio vary with the workload, making resource scaling and PD configuration interdependent. Exploiting this opportunity requires selecting effective configurations and realizing their benefits within transient resource-availability windows despite reconfiguration costs. We present PEARL, an asynchronous agentic RL system that coordinates external resource elasticity, temporary reuse of idle training GPUs, and adaptive PD execution. PEARL maintains a unified GPU--worker--role state and uses runtime profiles to predict rollout batch completion time, accounting for environment-induced reductions in decode concurrency. It selects the PD mode and ratio under the current GPU budget and translates each decision into an incremental transition plan that minimizes worker and role changes. Cost-aware switching and borrowing policies suppress transitions with insufficient expected benefit while ensuring timely return of training GPUs. Our evaluation show that PEARL achieves $2.17$--$2.79\times$ the throughput of fixed-resource ROLL across different LLMs. Compared with RLBoost+, throughput improves by up to approximately 26.9\% for Qwen3-8B and 36.3\% for Qwen3-30B-A3B.