Py-kvcache improves large language model caching performance with nvme ssds
Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs
Distributed, Parallel, and Cluster ComputingMachine Learning
Summary
Long conversations with artificial intelligence models can be slow because they need to remember everything said before. The authors studied how saving and reusing parts of these conversations, called key-value states, can speed up responses. They built a tool called py-kvcache that stores these states on fast NVMe disk drives and cleverly starts loading them before they're needed, making the process faster. They found that py-kvcache can speed up the time to start answering by up to twice compared to earlier systems, but using external cache depends on the setup and hardware.
What this means in practice
- •For llm deployment engineers: Speed up long-context large language model responses by integrating py-kvcache for efficient disk-based KV caching with asynchronous preloading.
- •For data center infrastructure teams: Optimize resource use by deciding when to enable external KV caching based on hardware and workload characteristics to reduce GPU memory pressure.
Authors
Joseph Kanichai, Tiziano De Matteis, Animesh Trivedi
Abstract
Prefix caching can reduce the time to first token (TTFT) of long-context LLM requests by reusing previously computed key-value (KV) states, but for short prefixes or fast GPUs, recomputation can be faster than loading from an external cache. We characterize this tradeoff in vLLM across GPU, CPU, and NVMe tiers using synthetic workloads, long-context benchmarks, production traces, and find that cache performance depends on transfer granularity, intermediate memory use, and when transfers enter the request schedule, not only on device bandwidth. These findings motivate py-kvcache, a vLLM KV Offload connector with asynchronous direct I/O, bounded shared staging, and scheduler-aware preloading, which starts disk reads while requests are still waiting, overlapping with compute. At 80k tokens, py-kvcache loading from disk is 2.0x faster than LMCache, with preloading contributing 1.34x. With GPU, CPU, and disk caching enabled, it is 1.23x faster than LMCache and within approximately 4% of the native vLLM KV Offload implementation. LongBench and SCBench show that these benefits extend to irregular prefix chains and multi-turn workloads. Bailian trace replays improve TTFT on a weaker GPU, but on an H100 the average request falls below the break-even point and GPU memory alone retains enough prefixes. External KV caching should therefore be treated as a setup specific admission decision. The py-kvcacheimplementation is available at: https://github.com/atlarge-research/py-kvcache.