TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts
2026-08-24 • Artificial Intelligence
Artificial IntelligenceMachine Learning
AI summaryⓘ
The authors focus on improving the speed of large language model (LLM) rollouts, where some requests take much longer to finish and slow down the whole process. They introduce TailSieve, a method that smartly routes these long requests separately and balances workload across computing replicas. TailSieve uses partial results from earlier runs to identify which requests might be slow, without needing extra training. This approach speeds up rollouts and helps specialized decoding methods work better, all while keeping the model's output quality consistent.
Large Language ModelsRolloutsLong-tail latencyRoutingReplica allocationSpeculative decodingOn-policy distillationConcurrencyThroughputLoad balancing
Authors
Tianqi Xu, Lu Lv, Haoyang Huang, Wenjie Huang, Zhanming Shen, Yuhao Shen, Baolin Zhang, Xinyi Hu, Shuang Ge, Jun Dai, Tianyu Liu, Suorong Yang, Zhikai Li, Ye Bai, Jun Zhang, Lei Chen, Yue Li, Mingchen Wan
Abstract
Large-scale rollouts have become a core component of modern LLM systems, spanning reinforcement learning (RL) post-training, on-policy distillation (OPD), and sampling-heavy evaluation pipelines. Unlike online serving, which is typically optimized for request-level latency and throughput, a small number of long-tail generations can dominate the end-to-end makespan of an entire rollout step. In practice, rollout requests are often routed uniformly across replicas, which can place extremely long generations inside high-concurrency decoding batches. To address this, we present TailSieve, a partial-rollout-guided framework that jointly controls tail routing and replica allocation for LLM rollouts. In an idealized setting with known completion lengths, we show that makespan-optimal routing in the long-tail regime combines tail isolation with load balancing, and that a simple top-k policy closely approximates this offline optimum. Leveraging the observation that long-tail prompts tend to remain long-tailed across policy updates, TailSieve uses partial rollouts as a training-free signal for identifying candidate tail groups. A hierarchical controller then jointly adapts the number of isolated groups and the replica split between the tail and bulk pools using collected response-work history and a measured concurrency-throughput model. TailSieve achieves up to 1.67x routing-only speedup over uniform group routing. The resulting low-concurrency tail pool further enables route-specialized speculative decoding with MTP or DFlash, achieving up to 2.59x speedup over uniform routing. Selected prompts are regenerated under the current policy, preserving on-policy generation and avoiding additional routing-induced length bias in steady state.