Papers for

cloud gpu infrastructure teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

TempoKV improves memory caching for faster large language model serving

TempoKV: Timely Staging of LLM KV Caches for Memory-Semantic Flash

Abstract: Reusable prefix key-value (KV) caches can outgrow GPU memory in large language model (LLM) serving. A memory-semantic flash hierarchy offers SSD-backed capacity with a limited fast tier, but a logical KV hit is not necessarily ready for GPU retrieval. Demand staging exposes SSD latency, whereas immediate staging can reserve fast-tier capacity long before retrieval begins. We present TempoKV, a timing-aware resource-commitment layer that separates early knowledge of reuse from the acquisition of staging resources. It records reusable-KV hits as metadata-only claims and requests commitment when the runtime-estimated time until retrieval falls to the storage-estimated time needed to make KV resident and protected against eviction. These estimates adapt to runtime progress and staging state, while commitment remains subject to available protected capacity. We implement TempoKV in vLLM and LMCache on an SSD-backed CXL memory device without changing request scheduling. Across two models and three prefix cache ratios, TempoKV reduces protected fast-tier byte-time per request by 63-91% versus immediate staging while retaining much of the serving benefit of advance staging. In a fast-tier capacity sweep, output throughput and p95 time to first token (TTFT) remain nearly unchanged as capacity decreases from 100 to 25 GiB. Compared with unmodified LMCache's Device-DAX L1 configuration, TempoKV reduces p95 TTFT by up to 48.0% and increases output throughput by up to 27.8%.

Mon 28 SeptDistributed, Parallel, and Cluster ComputingArtificial IntelligenceMachine Learning
The gist
Large language models save parts of their memory called key-value caches to speed up responses, but these caches can be too big for fast GPU memory. The authors designed TempoKV, which smarter manages when to move these caches into fast memory by predicting when they will be needed. This method saves memory resources and keeps the model running quickly even with less fast memory available. TempoKV was tested on real hardware and showed big improvements in speed and efficiency compared to previous methods.
Open → 2609.35065v1

Nested sequence parallelism speeds up training of long context large language models

NSP: Accelerating Variable-Length LLM Training via Nested Sequence Parallelism

Abstract: Long-context LLM training on long-tailed corpora faces a central communication--balance tradeoff. Such sequence-length heterogeneity makes any single sequence-parallelism (SP) degree a poor fit for the workload: a small degree leaves the few long sequences badly imbalanced, while a large degree forces the many short sequences that dominate the workload to pay excessive communication. Existing dynamic-SP systems mix SP degrees within a batch, but to run several groups at once they partition the GPUs into disjoint groups, which reintroduces imbalance across groups and forces costly micro-batch workarounds. We present NSP, a sequence-parallel training system that resolves this tradeoff by nesting differently sized SP groups on shared GPUs within a single training iteration. This lets long sequences use larger SP groups while keeping short sequences on smaller ones, so communication is incurred only where needed and load is balanced per GPU rather than per group. NSP realizes this idea with a tree-structured routing planner that assigns sequences under memory constraints and an executor that exploits the resulting hierarchy through inter-level phase streaming and tree-level recomputation. NSP supports common SP backends and requires no model changes. We evaluate NSP on Qwen3-MoE workloads with up to 384K-token contexts across multiple long-tail datasets on an internal production GPU cluster. Across these settings, NSP consistently improves end-to-end training throughput, outperforming Static SP by up to 1.48x and FlexSP by up to 1.16x.

Sat 19 SeptDistributed, Parallel, and Cluster Computing
The gist
Training large language models with very long and varied-length text sequences is slow because computers must balance work fairly while communicating enough data. The authors present NSP, a new method that groups sequences by length in a nested way on shared GPUs, so longer sequences get more resources and shorter ones get less, reducing wasted communication. This leads to faster overall training without changing the model itself. Tests on long-text datasets show NSP outperforms previous methods in training speed.
Open → 2609.22755v1