Dual-precision memory boosts large language model serving throughput

DPS: Dual-Mode Precision LLM Serving with Semi-Unified Memory

Distributed, Parallel, and Cluster ComputingArtificial IntelligenceEmerging TechnologiesPerformanceSoftware Engineering

Summary

Running large language models uses different kinds of memory for storing both the model and temporary data. This paper presents a system that changes how memory is used depending on workload: it normally runs the model at full accuracy, but when the system is busy, it switches to a simpler version to save memory and use that space for temporary data. By doing this, the system can handle more requests faster without losing much accuracy. The authors tested this approach and found it can more than double throughput while keeping accuracy close to usual standards.

What this means in practice

  • For llm deployment engineers: Serve large language models more efficiently by dynamically adjusting model precision to increase throughput during high-demand periods.
  • For cloud service operators: Improve server efficiency and meet latency guarantees by converting idle weight memory into cache during workload spikes.

Authors

Xuan Truong Nguyen, Tien Son Pham, Tuan Duc Chu, Wookeun Jung, Thanh Tuan Dao

Abstract

Existing LLM serving systems virtualize and optimize KV-cache memory, but treat model-weight memory as fixed throughout execution. Recent work on multi-precision model representations challenges this design by allowing a single stored model to support both full-accuracy and lower-precision execution, making the effective weight footprint runtime-dependent. This creates an opportunity under bursty workloads, where temporary spikes in KV-cache demand often determine throughput and SLO compliance. We present DPS, a dual-precision LLM serving system that turns weight memory into an elastic resource: under normal load, DPS serves the full-accuracy model; under KV pressure, it switches to a nested, lower-precision variant and repurposes unused weight memory for KV cache blocks. DPS is built on Semi-Unified Memory (SUM), which partitions the weight region into a persistent lower-precision sub-region and a shared region that alternates between residual weight tensors and KV-cache blocks, preserving compatibility with paged KV-cache management. We implement DPS on top of vLLM and evaluate it across both dense and MoE models and various production workload traces. Our results show that \sysname improves sustained throughput by $2.1$--$3.3\times$ and effective pass@1 by up to $+41$\,pp over Static FP16, while preserving FP16-class accuracy.