PackServe improves large language model request scheduling efficiency

PackServe: SLO-Aware Request Scheduling for Agentic LLM Serving at Scale

Distributed, Parallel, and Cluster Computing

Summary

Scheduling many requests for large language models (LLMs) is hard because it needs to reuse cached work, meet strict timing goals, and save GPU resources. The authors created PackServe, a scheduler that predicts processing delays and smartly groups requests on fewer GPUs while keeping speed and cache benefits. This reduces GPU usage significantly compared to previous methods, saving energy and cost without slowing down response times. PackServe is already used in a big production system with over 1000 GPUs.

What this means in practice

  • For cloud service operators: Reduce GPU resource consumption while maintaining strict latency guarantees for large language model serving at scale.
  • For data center infrastructure teams: Lower operational GPU hours and improve throughput by better request packing and cache reuse in multi-GPU clusters running LLM workloads.

Authors

Zhiyuan Tan, Dejiang Zhu, Jingzhe Jiang, Yihao Zheng, Yang Tian, Tao Wang, Minchen Yu

Abstract

Request scheduling is a key challenge in large-scale clusters serving agentic large language model (LLM) workloads. An effective scheduler must preserve key-value cache (KVC) reuse across long, shared prefixes, meet token-level latency service-level objectives (SLOs), and minimize GPU resource footprint. Existing schedulers struggle to reconcile these requirements: request consolidation can sacrifice cache locality and increase prefill/decode interference, compromising both SLO attainment and resource efficiency. We present PackServe, a scheduler designed to reduce resource costs while meeting latency SLOs for agentic LLM serving. PackServe uses compact white-box models to predict latency under prefill/decode interference. Guided by these predictions, it packs requests onto fewer serving instances while preserving KVC reuse and SLO constraints, trading available latency headroom for improved per-instance throughput. Evaluation on 64 NVIDIA H20 GPUs shows that PackServe uses up to 16.8% and 24.6% fewer GPU-hours than state-of-the-art schedulers under 30-ms and 50-ms TPOT targets, respectively, while meeting the target TPOT objectives. PackServe has also been deployed in our production cluster comprising over 1000 GPUs, where it reduces the resource footprint by 34.7% compared with the original production scheduler.