Spexis improves multi-GPU large language model inference efficiency

Spexis: Speculative Lookahead Scheduling for LLM Inference

Machine LearningDistributed, Parallel, and Cluster Computing

Summary

Running large language models (LLMs) on many GPUs efficiently is difficult because parts of the model need to work together and share data. The authors present Spexis, a system that guesses future computations in parallel without using extra memory, so it can speed up processing without slowing down. It also plans ahead to reduce wasted work and memory use. This helps make LLM applications faster across many GPU setups.

What this means in practice

  • For cloud ai service providers: Improve the speed and cost efficiency of hosting large language models on multi-GPU clusters using speculative parallelism.
  • For data center operators: Optimize GPU resource utilization and reduce memory bottlenecks during inference serving of large models.

Authors

Hyungyu Jung, Jaehyeok Yu, Hoonseo Choi, Sungkyun Kim, Jinho Lee, Jiwon Seo

Abstract

Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism. Rather than using speculative decoding only to accelerate token generation, Spexis runs speculation in parallel with normal execution, introducing a new parallelism axis without increasing KV-cache memory usage. This improves memory efficiency and helps mitigate the bottlenecks of multi-GPU inference. Spexis further uses lookahead scheduling to predict speculation quality and future memory pressure, allowing it to reduce wasted speculation, KV-cache eviction, and recomputation. Built on top of vLLM, Spexis largely improves serving performance across a range of GPU configurations, achieving speedups of up to 34% over a baseline that uses the optimal combination of pipeline and tensor parallelism. Spexis's source code is publicly available at https://github.com/mlsys-seo/spexis.