Papers for
data center infrastructure teams
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
QuantForge improves four-bit quantization for large language models
QuantForge: Discovering Residual Decompositions for MXFP4 Post-Training Quantization
Abstract: Four-bit post-training quantization can reduce the memory demands of large language models, but preserving accuracy under strict MXFP4 W4A4 requires coordinating several design choices. Coordinate transforms change block-encoding errors, which in turn affect the residuals propagated through the network. The useful algorithmic decomposition is therefore not fully known before search. LLM-driven program evolution offers a way to explore these choices, but performance scores alone do not explain which design should change next. We introduce QuantForge, a PTQ discovery system that records competing explanations, selects controls that distinguish them, and checks that successor code implements the resulting conclusions. This residual compilation guides program revisions while retaining useful programs even when their original explanations are rejected. Remeasuring the revised program reveals the next error to address. This process discovers HiRes, a fixed MXFP4 quantizer that shapes coordinates, refines legal code assignments, and recovers errors along attention and MLP paths. Each stage acts on residuals measured after the preceding stage has executed. Across seven tasks, HiRes achieves the lowest seven-model Robust Fit (0.09300) and the lowest quantized Fit-7 at 32B. In matched-budget comparisons of LLM-driven program evolution, each with 240 evaluator calls, QuantForge reaches a held-out transfer target in six of eight runs, compared with three each for textual memory and reflection memory, and one for score-only evolution, despite evaluating fewer new programs. These results show that QuantForge improves the discovery of transferable PTQ algorithms by turning controlled evidence into subsequent program changes.
PackServe improves large language model request scheduling efficiency
PackServe: SLO-Aware Request Scheduling for Agentic LLM Serving at Scale
Abstract: Request scheduling is a key challenge in large-scale clusters serving agentic large language model (LLM) workloads. An effective scheduler must preserve key-value cache (KVC) reuse across long, shared prefixes, meet token-level latency service-level objectives (SLOs), and minimize GPU resource footprint. Existing schedulers struggle to reconcile these requirements: request consolidation can sacrifice cache locality and increase prefill/decode interference, compromising both SLO attainment and resource efficiency. We present PackServe, a scheduler designed to reduce resource costs while meeting latency SLOs for agentic LLM serving. PackServe uses compact white-box models to predict latency under prefill/decode interference. Guided by these predictions, it packs requests onto fewer serving instances while preserving KVC reuse and SLO constraints, trading available latency headroom for improved per-instance throughput. Evaluation on 64 NVIDIA H20 GPUs shows that PackServe uses up to 16.8% and 24.6% fewer GPU-hours than state-of-the-art schedulers under 30-ms and 50-ms TPOT targets, respectively, while meeting the target TPOT objectives. PackServe has also been deployed in our production cluster comprising over 1000 GPUs, where it reduces the resource footprint by 34.7% compared with the original production scheduler.
Extender reduces memory needs for transformer attention with new channel
The Extender: A Log-Structured Transformer
Abstract: We introduce the Extender, a log-structured variant of the standard Transformer architecture. In a standard Transformer, each layer communicates with subsequent layers exclusively via the residual $\mathbf{h}$, a superposition channel. The Extender adds a concatenation channel $\mathbf{x}$: each layer $\ell$ emits both a residual update $δ_\ell$ which is added to $\mathbf{h}$, and a much smaller extension $ε_\ell$ which is appended to $\mathbf{x}$. While both the FFN and $\mathbf{q}$ see $\mathbf{h}$, the attention $\mathbf{kv}$ projections take only $\mathbf{x}$ as input. As a result, the fully extended $\mathbf{x}$ contains the complete input for the $\mathbf{kv}$ projections of all layers, reducing the persistent attention memory footprint from $2Ld_{model}$ to $\sum|ε_\ell|$. We find that with $|ε_\ell|=32$, the Extender matches Transformer accuracy on short-context (CORE) tasks at 199M-924M parameters, and exceeds Transformer accuracy on long-context (RULER) workloads, again at 924M parameters. For our 1664-wide, 924M model, the Extender's persistent attention memory footprint is $104\times$ smaller than MHA. The memory savings grow with model width.
Multi-asic switches boost network performance with circuit-switched indirection
The Power of Indirection: Scaling Switches Beyond Silicon Boundaries
Abstract: The slowdown of Moore's law and the area limit of monolithic integration have made chiplet-based designs inevitable across many domains, including network ASICs. However, combining multiple network ASICs together poses a fundamental challenge: maintaining sufficient inter-ASIC bandwidth to match the performance of an idealistic single-ASIC design. Providing full bandwidth is prohibitively expensive as it requires valuable forwarding capacity, while reducing inter-ASIC bandwidth creates severe performance bottlenecks. We propose a novel multi-ASIC switch architecture that introduces a circuit-switched indirection layer in front of the ASICs. This layer flexibly remaps ingress ports across ASICs, localizing traffic and minimizing inter-ASIC communication based on observed patterns. Our system, Fastroute, combines packet and circuit switching to deliver performance comparable to a single-ASIC switch while reducing inter-ASIC bandwidth requirements. This frees up capacity for external network interfaces, allowing Fastroute to outperform traditional non-oversubscribed multi-ASIC designs. Our hardware prototype demonstrates the system's functional feasibility by evaluating it on an LLM training workload. By reducing bandwidth and power overhead, Fastroute bridges the gap between silicon fabrication limits and soaring application demands. It provides an efficient transition to multi-ASIC switches, enabling bandwidth and radix demand to be met without waiting for the next ASIC generation.
Pattern matching improves online resource allocation in cloud networks
Pattern-Aware Virtual Network Embedding Optimization for Cloud Data Centers
Abstract: The network virtualization (NV) technology has enabled the sharing of multiple resources among virtual networks (VNs) in cloud data centers. One of the key challenges is to allocate resources in real-time for virtual network request (VNR), which is known as online virtual network embedding (VNE). However, the existing online VNE methods do not exploit the multi-dimensional complementary relationship among diverse VNRs, resulting in the fragmentation and waste of substrate resources. In this paper, we propose the pattern matching based online VNE approach by constructing appropriate matching rules among observed patterns to maximize resources utilization. We devise the clustering based VNRs quantization method and conduct rigorous study on the pattern combination filtering problem. Then, we utilize the column generation to solve it and construct the pattern matching rules. Based on the rules, we propose an online pattern matching VNE algorithm with linear worst-case complexity. Evaluation on a 106-server testbed using Alibaba production cluster trace dataset shows that our algorithm achieves close-to-offline performance and more accepted workloads that outperforms traditional designs by 25%-30%.
Fast method commits many power generators under grid limits
Fast Relax-and-Round Unit Commitment with Topological Constraints
Abstract: Recent developments in the US knowledge economy have created a significant growth in datacenter loads, with two major consequences for power generation. First, to compensate for growth in load, datacenters are encouraged to bring their own generating units. Second, in search of the remaining pools of dispatchable generation, utilities are increasingly turning to subtransmission and distribution level generating assets. Coupled with increasing loads and the resulting tighter grid conditions, both trends are likely to create a need to commit a large number of localized generating units under grid constraints. We propose an extension of Relax-and- Round Unit Commitment (RRUC) that is capable of committing generating units for larger problems faster than conventional methods, while staying within intertemporal and spatial MVA constraints. We demonstrate the performance of RRUC using synthetic congestible test systems ranging from 100 to 20,000 buses. RRUC consistently finds low cost solutions, independent of the problem size, and its run time increases sub-quadratically in the number of buses. RRUC can solve the 100 bus system in less than a second and the 20,000 bus system 7 minutes. In contrast, a leading state of the art solver cannot find a feasible solution to the 100 bus system in 15 minutes.
Py-kvcache improves large language model caching performance with nvme ssds
Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs
Abstract: Prefix caching can reduce the time to first token (TTFT) of long-context LLM requests by reusing previously computed key-value (KV) states, but for short prefixes or fast GPUs, recomputation can be faster than loading from an external cache. We characterize this tradeoff in vLLM across GPU, CPU, and NVMe tiers using synthetic workloads, long-context benchmarks, production traces, and find that cache performance depends on transfer granularity, intermediate memory use, and when transfers enter the request schedule, not only on device bandwidth. These findings motivate py-kvcache, a vLLM KV Offload connector with asynchronous direct I/O, bounded shared staging, and scheduler-aware preloading, which starts disk reads while requests are still waiting, overlapping with compute. At 80k tokens, py-kvcache loading from disk is 2.0x faster than LMCache, with preloading contributing 1.34x. With GPU, CPU, and disk caching enabled, it is 1.23x faster than LMCache and within approximately 4% of the native vLLM KV Offload implementation. LongBench and SCBench show that these benefits extend to irregular prefix chains and multi-turn workloads. Bailian trace replays improve TTFT on a weaker GPU, but on an H100 the average request falls below the break-even point and GPU memory alone retains enough prefixes. External KV caching should therefore be treated as a setup specific admission decision. The py-kvcacheimplementation is available at: https://github.com/atlarge-research/py-kvcache.