Papers for

llm deployment engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Kv cache compression method retains long and off context query strength

Cartridges++: KV Cache Compression without Off-Context Derailment

Abstract: Serving long documents to a Large Language Model (LLM) repeatedly is expensive: computations grow with context length, and the memory footprint of the key-value (KV) cache balloons. Compressed KV (CKV) representations aim to mimic the cache of a document and are typically computed once and for all, ahead of inference time. Methods to obtain CKVs range from drop mechanisms that reduce their number of columns, to learned approaches. Among the latter, Cartridges have emerged as a leading compression method, learning compact KV representations through distillation on relevant Q/A pairs. While existing evaluations focus primarily on whether Cartridges and other CKVs yield approximately similar responses to document-related, on-context queries, we investigate the crucial deployment question of whether they can handle off-context queries, something the native KV representation is particularly good at, thanks to the mechanics of attention. We observe a fundamental trade-off: while Cartridges perform better for on-context queries, heuristic-variants preserve better the original LLM's ability to operate off-context. We measure this through their capability to avoid context contamination in their response, retain general knowledge, and follow instructions. We propose Cartridges++, simple modifications to cartridges that retain off-context abilities at small or negligible cost. The router variant decides at inference time whether the query should use the learned long-context memory, while the data-mixing variant allocates a small fraction of training Q/As to queries outside the reference long document. Our study shows that assessing CKVs on document utility alone can mask substantial degradation in broader model capabilities, yet those issues can be fixed with benign changes to CKV inference or training.

Mon 28 SeptMachine Learning
The gist
Working with long documents on large language models uses a lot of computer memory and time, so people try to compress the memory it needs. The authors study a popular compression method called Cartridges, which works well for questions about the original document but struggles when questions are unrelated to that document. They find a trade-off between remembering document details and handling general queries. They propose Cartridges++, which modify the compression to keep the benefits for document queries while also preserving the model’s ability to answer off-topic questions well.
Open → 2609.35621v1

Dual-precision memory boosts large language model serving throughput

DPS: Dual-Mode Precision LLM Serving with Semi-Unified Memory

Abstract: Existing LLM serving systems virtualize and optimize KV-cache memory, but treat model-weight memory as fixed throughout execution. Recent work on multi-precision model representations challenges this design by allowing a single stored model to support both full-accuracy and lower-precision execution, making the effective weight footprint runtime-dependent. This creates an opportunity under bursty workloads, where temporary spikes in KV-cache demand often determine throughput and SLO compliance. We present DPS, a dual-precision LLM serving system that turns weight memory into an elastic resource: under normal load, DPS serves the full-accuracy model; under KV pressure, it switches to a nested, lower-precision variant and repurposes unused weight memory for KV cache blocks. DPS is built on Semi-Unified Memory (SUM), which partitions the weight region into a persistent lower-precision sub-region and a shared region that alternates between residual weight tensors and KV-cache blocks, preserving compatibility with paged KV-cache management. We implement DPS on top of vLLM and evaluate it across both dense and MoE models and various production workload traces. Our results show that \sysname improves sustained throughput by $2.1$--$3.3\times$ and effective pass@1 by up to $+41$\,pp over Static FP16, while preserving FP16-class accuracy.

Mon 28 SeptDistributed, Parallel, and Cluster ComputingArtificial IntelligenceEmerging Technologies
The gist
Running large language models uses different kinds of memory for storing both the model and temporary data. This paper presents a system that changes how memory is used depending on workload: it normally runs the model at full accuracy, but when the system is busy, it switches to a simpler version to save memory and use that space for temporary data. By doing this, the system can handle more requests faster without losing much accuracy. The authors tested this approach and found it can more than double throughput while keeping accuracy close to usual standards.
Open → 2609.34380v1

Py-kvcache improves large language model caching performance with nvme ssds

Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs

Abstract: Prefix caching can reduce the time to first token (TTFT) of long-context LLM requests by reusing previously computed key-value (KV) states, but for short prefixes or fast GPUs, recomputation can be faster than loading from an external cache. We characterize this tradeoff in vLLM across GPU, CPU, and NVMe tiers using synthetic workloads, long-context benchmarks, production traces, and find that cache performance depends on transfer granularity, intermediate memory use, and when transfers enter the request schedule, not only on device bandwidth. These findings motivate py-kvcache, a vLLM KV Offload connector with asynchronous direct I/O, bounded shared staging, and scheduler-aware preloading, which starts disk reads while requests are still waiting, overlapping with compute. At 80k tokens, py-kvcache loading from disk is 2.0x faster than LMCache, with preloading contributing 1.34x. With GPU, CPU, and disk caching enabled, it is 1.23x faster than LMCache and within approximately 4% of the native vLLM KV Offload implementation. LongBench and SCBench show that these benefits extend to irregular prefix chains and multi-turn workloads. Bailian trace replays improve TTFT on a weaker GPU, but on an H100 the average request falls below the break-even point and GPU memory alone retains enough prefixes. External KV caching should therefore be treated as a setup specific admission decision. The py-kvcacheimplementation is available at: https://github.com/atlarge-research/py-kvcache.

Thu 10 SeptDistributed, Parallel, and Cluster ComputingMachine Learning
The gist
Long conversations with artificial intelligence models can be slow because they need to remember everything said before. The authors studied how saving and reusing parts of these conversations, called key-value states, can speed up responses. They built a tool called py-kvcache that stores these states on fast NVMe disk drives and cleverly starts loading them before they're needed, making the process faster. They found that py-kvcache can speed up the time to start answering by up to twice compared to earlier systems, but using external cache depends on the setup and hardware.
Open → 2609.11744v1