Papers for

large scale ai training teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

HyDra improves training speed and balance for long AI sequences

HyDra: Demystifying and Taming Dynamic Context Parallelism at Production Scale

Abstract: Long-context training runs on sequences whose lengths span orders of magnitude, and dynamic context parallelism (DCP) gives each sequence its own CP degree. Existing DCP systems either do not scale or perform poorly on mainstream models, leaving Megatron-Core (Mcore) DCP as the only option at production scale. Mcore DCP, however, sizes each degree to fit memory, which grows linearly with length while attention grows quadratically, so comparable token counts hide unequal computation. In our production 256K-context training job on more than 11K GPUs under Mcore DCP, per-rank microbatch times differ by up to 5x at similar token counts. The skew leads to a 46% pipeline bubble and a 13% data-parallel bubble. We present HyDra, a scalable load-driven DCP system. Its scheduler balances computation by pulling every rank toward one load target, and balance in turn makes that target solvable in closed form. It places sequences with lazy heaps rather than whole-pool scans. That balance asks for CP degrees larger than memory requires, so its CP engine nests an inner Ulysses group in a shallow outer ring, letting a higher degree lower computation and communication together. Evaluation at both scales shows consistent gains. On a 512-GPU testbed, HyDra raises throughput over Mcore DCP by 1.18x on average at 32K context and 2.48x at 256K. On a 2,048-GPU production job, it shrinks the pipeline bubble from 36% to 14%, cuts scheduling time by 2.3x, and raises throughput by 1.10-1.43x (avg. 1.25x) over Mcore DCP and 1.33-1.90x (avg. 1.59x) over static CP.

Mon 28 SeptDistributed, Parallel, and Cluster Computing
The gist
Training AI models on very long sequences can take uneven time because different parts handle different amounts of work. The researchers found that a common way to split this work caused big slowdowns and inefficiencies. They built HyDra, a system that smartly balances the work across all computers running the training so everything finishes more evenly. This new system made training faster and smoother, especially for very long sequences.
Open → 2609.34318v1