Data center flexibility varies with time duration reliability and cluster size

Beyond Scalar Flexibility: From Eligible AI Workloads to Dependable Load Relief

Distributed, Parallel, and Cluster Computing

Summary

Power grids often assume that data centers can reduce their electricity use by a fixed percentage when needed, but actual data showing how much can be cut and for how long is lacking. The authors analyzed a large set of GPU power use data from many clusters to understand how measurable and dependable this flexibility really is over various time spans. They found that flexibility decreases as the duration increases and that combining multiple clusters helps but has limits due to correlations in power use. Their findings offer a more detailed way to specify flexibility in power contracts instead of using simple fixed percentages.

data center flexibilitypower demand responseGPU workloadload curtailmentMonte Carlo simulationduration dependencecross-cluster covariancedemand response contractsavailability reliability

Authors

Meiyi Li

Abstract

Grid studies often represent data-center flexibility as a fixed percentage of load, although no public production trace has shown how much eligible load persists across event durations or co-moves across clusters. We reconstruct 4,439 hourly power observations from a 185-day trace of 155,410 GPUs and derive a workload-semantic flexibility envelope. The fleet's time-averaged Monte Carlo median facility demand is 55.8 MW, while immediate eligible curtailment averages 3.55 MW after retaining allocated-GPU idle power: 12.1% of workload power and 6.35% of median facility power. Under full realization of that eligibility, 95%-available relief falls from 2.51 MW for one hour to 2.32 MW for four hours and 1.95 MW for 24 hours; a common realizable fraction q scales every value exactly by q. A mean-calibrated scalar overstates these quantities by 17%, 25%, and 47%, while a scalar tail-calibrated at four hours understates the one-hour product by 6% and overstates the 24-hour product by 17%; the share that reproduces the surface varies by a factor of 1.6 across durations and reliability levels. Aggregating 13 clusters raises four-hour firmness from 0.38 to 0.66, but cross-cluster covariance limits the gain. The production scheduler exposes almost no additional delay-based capacity: newly deferrable arrivals average 0.008 MW and have zero 95%-available capacity. These results replace an assumed flexibility percentage with duration, reliability, portfolio, and realizability terms that can be written into interconnection and demand-response contracts.