Summary
AI data centers use a huge amount of electricity and face limits based on how much power the electric grid can supply. The authors explain that to keep these centers running well, their power systems—from the electric grid connection down to the computer chips—need to work together and handle quick changes in power supply. They talk about evolving technologies and control methods that help keep AI data centers stable and flexible, like better cooling, energy storage, and smarter workload management. Their main idea is that designing the entire power and computing system as one connected unit is necessary for future AI infrastructure.
What this means in practice
- •For data center engineers: Design power systems that manage rack, facility, and grid stability for AI hardware under varying loads and grid conditions.
- •For power grid operators: Develop grid-connection policies and control strategies that accommodate AI data centers as flexible, grid-interactive loads to improve overall stability.
A survey. It maps existing work.
Authors
Yubo Song, Rui Kong, Takuro Umihara, Pooya Davari, Frede Blaabjerg, Subham Sahoo
Abstract
The rapid growth of artificial intelligence (AI) computing is transforming data centers into large, dynamic electrical loads. Their deployment is primarily constrained by energy availability and grid-connection capacity, which is further aggravated by the ability of power-delivery architectures, control systems, and computing workloads to operate reliably during fast grid disturbances. This article presents a technological perspective on AI data centers as grid-interactive computing systems. First, it reviews grid-integration bottlenecks, evolving connection policies, grid-code requirements, which has fostered new technological trends via spatio-temporal flexibility available through workload orchestration, cooling systems, on-site resources, and energy storage. Second, it maps the evolution of power-delivery architectures from medium-voltage grid interfaces to chip-level, discussing higher-voltage DC distribution, solid-state transformers, wide-bandgap devices, advanced chip-level power delivery, and liquid cooling. Third, it establishes a three-level stability framework spanning rack-level DC-bus dynamics, facility-level converter interactions, and system-level grid-coupled behavior. The framework connects dominant instability mechanisms, including constant power load effects, impedance interactions, forced oscillations, and operating-mode transitions, with suitable modeling, assessment, and mitigation approaches. Synthesizing these topics, this article highlights grid-to-chip co-design as a central requirement for scalable AI infrastructure, linking computing workloads, power-delivery systems, energy buffers, and grid operation.