How words turn into network data during large AI model training

The Life of a Token: from Words to Bits on the Wire

Distributed, Parallel, and Cluster ComputingMachine LearningNetworking and Internet ArchitecturePerformance

Summary

Training huge AI language models involves turning words into bits that travel across many computers working together. The authors explain the detailed process of how words are broken down, turned into numbers, and then sent as data across networks in high-performance computing systems. They use examples from a famous book to show how different parts of the model and training setup affect the communication between computers. This helps people understand the network needs for training large language models.

What this means in practice

  • For data center network engineers: Design network infrastructure with proper capacity and timing for training massive language models efficiently.
  • For ai system architects: Plan parallelization and communication strategies tailored to the communication patterns of large language model training.

A survey. It maps existing work.

Authors

Davide Avesani, Pengwenlong Gu, Sotiris Skaperas, Stefano Secci

Abstract

Large Language Models (LLMs) transform vast collections of unstructured text into semantic patterns used for language generation and reasoning tasks. Behind their ease of use lies a complex process: words become tokens, tokens become vectors, and vectors ultimately give rise to streams of bits that flow through High-Performance Computing (HPC) systems. As modern LLMs grow to billions or trillions of parameters, this path increasingly unfolds across thousands of interconnected accelerators, making the underlying communication fabric a critical and often opaque component of model training. This tutorial aims to walk the reader through the journey from words to network traffic, shedding light on how language is translated into communication flows within HPC training systems. Using concrete examples from Dante's Divine Comedy, we illustrate how model architecture, tokenization, embeddings, and parallelization strategies shape the volume, structure, and timing of data exchanged across the network. We combine architectural analysis with analytical traffic models and numerical examples to characterize the communication requirements of LLM training. We try to demystify how words travel across the network and provide practical insights into the network requirements needed to support the journey from text to trained model.