Papers for

infrastructure architects

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

MegaGraph speeds up training on large-scale graph transformers efficiently

MegaGraph: Towards Efficient Training of Large-Scale Graph Transformers with Automated Hybrid Parallelism

Abstract: Graph Transformers (GTs) offer superior representation capabilities by overcoming the depth limitations and over-smoothing issues of traditional Graph Neural Networks (GNNs). However, scaling GTs to large graphs poses critical bottlenecks. Specifically, the attention score matrix and its associated topology-aware bias matrix jointly incur significant per-layer memory overhead, and heavy graph embedding layers result in severe workload imbalances. These characteristics are unique to GT training and are not addressed by parallelism techniques designed for either conventional GNNs or Transformers, making a dedicated solution necessary. This paper introduces MegaGraph, the first automated hybrid parallelism framework designed for efficient GT training. MegaGraph designs three specialized strategies, namely graph-aware context parallelism, heterogeneous pipeline parallelism, and hybrid data parallelism, to support efficient training on large-scale graphs. However, coordinating these three parallelism strategies yields an exponentially large configuration space. To address this complexity, an automatic search engine leverages precise cost models via a Profile - Model - Search workflow to identify the optimal parallelism configuration. Evaluations demonstrate that MegaGraph enables training on large-scale graphs where state-of-the-art baselines fail due to out-of-memory (OOM) errors. The framework reduces per-device peak memory by up to 77.8\% and achieves up to 4.51$\times$ training speedup while maintaining model accuracy.

Mon 28 SeptMachine Learning
The gist
Training special AI models called Graph Transformers on very large graphs is hard because they use a lot of memory and have uneven workloads. The authors created MegaGraph, a system that combines three smart ways to split up the work across many devices. To find the best way to do this efficiently, MegaGraph automatically searches through possible setups using careful cost models. This system lets training happen on big graphs where previous methods ran out of memory, making it faster and more memory-friendly without losing accuracy.
Open → 2609.34420v1