Papers for

data center network engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

VarioPath speeds up GPU cluster data sharing by managing PCIe traffic

VarioPath: Workload-Aware All-to-All Communication for PCIe GPU Clusters

Abstract: AlltoAllv communication is a critical primitive in distributed large-model inference, particularly for mixture-of-experts (MoE) models. The growing adoption of PCIe GPU systems for cost-efficient inference makes AlltoAllv performance on these systems increasingly important. Without a dedicated scale-up interconnect (e.g., NVLink or Infinity Fabric), PCIe GPU systems carry both intra-node and inter-node traffic through the PCIe hierarchy, where concurrent transfers can contend for PCIe link bandwidth. This link contention, compounded by skewed traffic distributions and dynamic traffic demand, makes efficient AlltoAllv scheduling challenging. Existing approaches are either poorly suited to PCIe GPU systems or incur substantial schedule synthesis overhead that reduces their practicality in real-world deployments. We present VarioPath, an efficient AlltoAllv scheduling framework for PCIe GPU systems. It combines an offline topology-aware analyzer with an online demand-aware scheduler. The analyzer records contention-free transfer patterns as AlltoAllv channels and exploits topology symmetry to build a compact catalog for efficient search. The online scheduler decomposes each AlltoAllv invocation's demand across a sequence of channels, adapting to rapidly changing and skewed traffic while incurring low planning overhead. Evaluation on four platforms (up to 256 GPUs) shows average AlltoAllv speedups of 5.88x over FAST and 1.72x over DeepEP. End-to-end experiments show that VarioPath reduces Qwen3 inference latency by up to 27.2% and Wan2.1 generation latency by 6.1%.

Mon 28 SeptDistributed, Parallel, and Cluster ComputingNetworking and Internet Architecture
The gist
Sharing information quickly between many GPUs is important for running large AI language models smoothly. On systems using PCIe connections, the data transfer can get jammed because many transfers compete for the same routes. The authors designed VarioPath, a method that smartly plans these data transfers based on the computer setup and current demands. This approach helps reduce delays and speeds up tasks like AI model inference on GPU clusters.
Open → 2609.34340v1

Flux schedules optical switches to speed up AI model training

Flux: Optimal Scheduling of Optical Circuit Switches for LLM Training

Abstract: Optical Circuit Switching (OCS) offers high bandwidth density and energy efficiency for LLM training, but incurs a non-negligible reconfiguration delay. Prior work typically schedules optical circuit switches independently of compute, using aggregate traffic demand to determine which circuits to provision and when. We argue that this separation creates a fundamental inefficiency: reconfigurations that ignore the compute timeline can stall communication, resulting in low circuit utilization and large buffer requirements. In this paper, we present Flux, a scheduler that optimally schedules optical circuit switches based on the structure of the entire workload. Flux remains effective across a wide range of switching speeds by reusing circuits and amortizing reconfiguration delay behind compute and communication. We show that Flux reduces training iteration time by up to $10\times$ and peak NIC buffer requirements by more than three orders of magnitude compared to traditional periodic schedulers.

Tue 22 SeptNetworking and Internet ArchitectureDistributed, Parallel, and Cluster Computing
The gist
Training large AI models needs fast and efficient data movement between computers. Optical circuit switches offer great speed but take some time to change connections, causing delays. The authors argue that planning switch changes without considering the training steps slows things down and wastes resources. They created Flux, a scheduler that plans switch changes based on the whole training work, reducing delays and buffer needs. Flux can make training up to ten times faster and greatly reduce memory use compared to older methods.
Open → 2609.25949v1

Fpga design tracks packets fast for reliable wide area networks

Scalable Packet Tracking on FPGAs for Erasure-Coded RDMA over Lossy WANs

Abstract: Modern AI workloads increasingly rely on scale across architectures that interconnect multiple datacenters to form a single "AI factory", overcoming the power and cooling constraints of individual sites. However, extending Remote Direct Memory Access (RDMA) across wide area networks (WANs) introduces fundamental challenges: multi-path packet reordering, high latency, and packet loss that severely degrade performance. While erasure coding (EC) has emerged as a promising mechanism for loss recovery, its effectiveness critically depends on efficient packet arrival tracking implemented in hardware. We present COmpact Multi-path Erasure-coded Tracking (COMET), the first fully hardware-offloaded packet-arrival tracking design implemented on an FPGA-based network interface card (NIC) for multi-path RDMA over lossy WANs. COMET employs a scalable cache-based architecture that supports operation at high link rates. Our evaluation shows that COMET sustains line rate operation at 400 Gbps and beyond. Critically, COMET decouples on-chip memory footprint from link Bandwidth-Delay Product (BDP), and its cache-based architecture (COMET Cache) enables supporting 6 times more concurrent connections than state-of-the-art (SOTA) SoC-based designs. These results demonstrate that scalable, fully hardware-offloaded packet-arrival tracking is practical on FPGA-based NICs at current data rates, and its architectural scalability extends to emerging 1.6 Tbps NICs and beyond.

Fri 18 SeptHardware Architecture
The gist
Connecting multiple data centers to work together for AI tasks is tricky because data sent over long distances can get mixed up, delayed, or lost. The authors show that a special computer chip called an FPGA can help keep track of data packets quickly and efficiently, even when some are lost along the way. Their design, called COMET, can handle very fast data connections without slowing down, making it easier to send data reliably over large networks. This can help data centers work better together at huge scales.
Open → 2609.21774v1

How words turn into network data during large AI model training

The Life of a Token: from Words to Bits on the Wire

Abstract: Large Language Models (LLMs) transform vast collections of unstructured text into semantic patterns used for language generation and reasoning tasks. Behind their ease of use lies a complex process: words become tokens, tokens become vectors, and vectors ultimately give rise to streams of bits that flow through High-Performance Computing (HPC) systems. As modern LLMs grow to billions or trillions of parameters, this path increasingly unfolds across thousands of interconnected accelerators, making the underlying communication fabric a critical and often opaque component of model training. This tutorial aims to walk the reader through the journey from words to network traffic, shedding light on how language is translated into communication flows within HPC training systems. Using concrete examples from Dante's Divine Comedy, we illustrate how model architecture, tokenization, embeddings, and parallelization strategies shape the volume, structure, and timing of data exchanged across the network. We combine architectural analysis with analytical traffic models and numerical examples to characterize the communication requirements of LLM training. We try to demystify how words travel across the network and provide practical insights into the network requirements needed to support the journey from text to trained model.

Thu 17 SeptDistributed, Parallel, and Cluster ComputingMachine LearningNetworking and Internet Architecture
The gist
Training huge AI language models involves turning words into bits that travel across many computers working together. The authors explain the detailed process of how words are broken down, turned into numbers, and then sent as data across networks in high-performance computing systems. They use examples from a famous book to show how different parts of the model and training setup affect the communication between computers. This helps people understand the network needs for training large language models.
Open → 2609.19924v1

Low-latency flow-based intrusion detection for FPGA smart network cards

FSNIC: A Low-Latency Flow-Based Intrusion Detection Architecture for FPGA SmartNICs

Abstract: Modern data centres require high-performance networking alongside effective real-time security. Traditional Intrusion Detection Systems (IDS) commonly rely on general-purpose processors and often struggle to inspect high-speed traffic at line rate without introducing latency or performance bottlenecks. Smart Network Interface Cards (NICs) provide an alternative by enabling computation directly within the network data plane. This work presents a machine learning-based IDS implemented within an FPGA-based SmartNIC pipeline. The system integrates P4-based packet parsing with a LogicNets IDS model implemented in RTL, enabling deterministic, low-latency inference. Compared with traditional stateless packet-level classifiers, the proposed stateful flow-based IDS introduces minimal state by aggregating features across packets, capturing behavioural patterns not observable at the packet level. Experimental results on the UNSW-NB15 dataset show that the flow-based IDS improves detection accuracy from 86.92\% to 97.68\% compared with stateless packet-level classification. We also evaluate the proposed IDS on CICIDS2017 and compare its real-time hardware performance with prior FPGA-based IDS designs. Through hardware-software co-design, the proposed IDS achieves 6~ns inference latency using only 846 LUTs, with no BRAM or DSP usage, demonstrating a low latency and resource efficient implementation.

Mon 14 SeptHardware Architecture
The gist
Data centers need strong and fast security systems to detect network attacks without slowing down traffic. The authors designed a machine learning system that runs directly on smart network cards with FPGAs, allowing quick and efficient recognition of attacks by looking at sequences of packets rather than single packets alone. This approach improves the accuracy of detecting threats and processes data with very low delay and minimal hardware resources. Their tests showed better detection rates and faster performance compared to earlier methods.
Open → 2609.16363v1

Optical switch scheduling boosts large language model training speed

Breaking the Duplex Barrier: Lane-Granularity OCS Scheduling for LLM Training

Abstract: Optical circuit switch (OCS) can reconfigure physical connectivity to match the predictable communication schedules of large language model (LLM) training. Although each OCS light path is physically simplex, existing demand-aware OCS schedulers allocate capacity in duplex-port pairs, forcing equal bandwidth in both directions and stranding capacity under asymmetric node-pair traffic. This paper present LACE, the first offline OCS schedule compiler that independently allocates transmit (TX) and receive (RX) lanes for LLM training. Without changing the selected collective algorithms, operation order, or rank placement, LACE reconstructs directed node-level demand, jointly determines which consecutive operations share a configuration and how many simplex circuits serve each direction, and realizes these allocations as physical lane bindings and optical paths under per-node lane-inventory and multi-OCS fabric constraints. Software acknowledgments carry feedback over independently provisioned return paths, while coordinated link configuration and recovery verify each configuration before communication resumes. On a separate three-server testbed using fixed topologies and matched per-port rate limits, LACE's asymmetric connectivity achieves $1.80\times$ speedup for communication replay and $1.27\times$ for GPT-2 training over a symmetric-topology baseline. At larger scale, simulations of LLaMA-3.1 70B and 405B schedules with sixteen 400-Gb/s ports per server show that LACE achieves $1.21$--$2.04\times$ communication speedup over the latest duplex OCS scheduler.

Sun 13 SeptNetworking and Internet Architecture
The gist
Large language models need lots of data to be sent between computers, but current optical switches limit data flow by treating connections as two-way equal paths. The paper introduces LACE, a system that breaks this limit by allowing one-way (simplex) data lanes, adapting to uneven communication needs. This approach speeds up data transfer during training without changing the overall training setup. Tests show it can make training faster by up to nearly twice as much in some cases.
Open → 2609.14253v1