Papers for
power system operators
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Powerbench evaluates language models on power system data tasks
PowerBench: A Benchmark for Agentic Retrieval and Reasoning in Power Systems
Abstract: Large language model (LLM) agents offer new opportunities for automated analysis in industry. However, rigorous evaluation of such agents-for example, within power system scenarios-remains hindered: real operational data are confidential, and existing public resources fail to fully capture the chained dependencies and heterogeneous evidence. To address this gap, we propose PowerBench, comprising (1) a generation framework that derives interconnected heterogeneous operational data through a common dependency chain, and (2) a synthetic dataset generated by this framework. The dataset covers 761 devices across 100 device types, with 13.35 million hourly telemetry records spanning two years and 24,939 operational documents. Building on this dataset, we construct 300 questions across three task families that evaluate frontier LLMs' ability to complete analysis tasks that require autonomous evidence retrieval and reasoning across interconnected and heterogeneous data under restricted tool calls and time budgets. Results demonstrate that the evaluated frontier LLMs remain challenged on these tasks: the best model reaches only 74.2% joint accuracy. Our trace analysis further reveals that model performance varies across evidence discovery, content retrieval, tool use, reasoning over evidence, and answer submission. These findings provide detailed insights for evaluating LLM agents and guiding their reliable deployment in industry. The framework, dataset, and benchmark tasks are available at https://github.com/open-compass/PowerBench.
NGBoost improves power demand forecasting with uncertainty estimates
Probabilistic electrical power demand forecasting with uncertainty quantification
Abstract: The majority of research on electricity consumption forecasting has focused on deterministic approaches, which generate a single point estimate for each time step in the forecasting horizon. However, the increasing penetration of renewable energy sources and the growing complexity of modern smart grids have introduced greater variability and uncertainty into power-system demand and operation. Consequently, probabilistic forecasting, which quantifies the uncertainty and variability associated with future electricity demand, is becoming increasingly important for reliable power-system planning and operation. This study presents an empirical comparison of four contemporary probabilistic forecasting models for electricity consumption, highlighting their respective strengths and limitations. We have performed comparision on real-world power systems related datasets. Across all power-consumption zones, NGBoost demonstrates superior probabilistic forecasting performance, achieving the lowest MAE and RMSE while providing well-calibrated uncertainty estimates with high prediction-interval coverage and reasonably narrow intervals. These results indicate that NGBoost offers a more accurate and reliable forecasting framework than Bayesian, Monte Carlo (MC) Dropout, and Gaussian Process Regression (GPR) models for the considered electricity consumption data.
Power grid outage analysis improved using single deep neural network
AC Power Flow Contingency Analysis Using a Single Deep Neural Network
Abstract: Contingency analysis using the AC power flow (AC-PF) model is a critical tool for accurate grid security assessment, but its computational burden increases with the number of operating scenarios and outage configurations to evaluate. Recent ML-based approaches typically require outage-specific training data, leading to offline training costs that scale with the number of contingencies. This work proposes a framework that reuses a single ML model trained solely on basecase AC-PF data to estimate post-contingency operating states under arbitrary single-line outages. The proposed approach formulates post-contingency state prediction as a fixed-point iteration. If the ML model is a deep neural network (DNN), we derive sufficient conditions that guarantee convergence and develop semidefinite programming (SDP) formulations to certify these conditions for a given DNN. Numerical tests on the IEEE 118-bus system demonstrate that the proposed SDP formulations are tight, that the certified conditions hold for all tested contingencies, and that the resulting method produces accurate post-contingency state estimates within only a few iterations.
Finite sample safety check for ai control of energy devices
Finite-Sample Probabilistic Safety Certification for AI-Based Grid-Edge Coordination
Abstract: Coordinating large population of flexible grid-edge devices can alleviate the need for time-consuming and capital-intensive network upgrades, and AI-based control methods such as multi-agent reinforcement learning or imitation learning are promising in their real-time decision scalability. However, system operators still need an independent and rigorous way to decide whether a given AI system is safe enough for deployment. This paper develops a finite-sample probabilistic safety certification framework for black-box AI decision models in closed-loop grid operation. The central idea is to reduce the complete input--AI--grid evaluator workflow to a binary unsafe outcome under an operator-defined safety specification, and then use exact binomial inference to certify the corresponding unsafe operation probability. Given a set of held-out calibration scenarios, the framework returns the tightest one-sided upper certificate and an accept/reject deployment criterion that controls the probability of false safety certification. Because the certification is for the calibration distribution that may deviate from the future operation, we further combine the nominal certificate with physically interpretable sample-space adversarial attacks, a concept widely used in AI to investigate the fragility of AI models. Case studies on grid-edge flexibility coordination with 1{,}000-agent AI models (independent parameters) verify the finite-sample safety guarantee and the value of integrating adversarial attacks into a rolling-window training-certification-deployment flow.
Power flow model includes droop control effects for better grid analysis
Droop-Aware Foundation Model Power Flow
Abstract: This paper develops a droop-aware extension of the GridFM power systems foundation model, embedding droop gains and frequency/voltage deadband parameters as per-bus node features to enable control-aware AC power-flow analysis. Existing power-flow datasets encode only static electrical features, conflating operating points from qualitatively different control regimes; this work resolves that gap by exposing droop and deadband parameters as structured node features, with deadband discontinuities handled through a smooth tanh approximation that preserves solver differentiability. A transformer-based graph neural network is pre-trained on masked reconstruction and fine-tuned on the resulting control-aware datasets. The framework is validated against PSCAD electromagnetic-transient simulations on a two-bus system (0.11% maximum steady-state error) and cross-validated against an independent PyPower droop solver on the IEEE 24-bus RTS. On the 24-bus system the surrogate attains R2 = 0.9996 for active generation and 0.0015 p.u. voltage-magnitude RMSE; scalability is confirmed on the IEEE 300-bus system (0.0036 p.u. RMSE, R2 = 0.9841 for voltage magnitude across 299,700 predictions). A three-mode control study further shows that the deadband widens the control-error distribution while leaving total droop compensation unchanged, establishing deadband width as an actionable node-level design feature.
Single graph neural network solves multiple power system analysis tasks
Unified Heterogeneous Graph Neural Network solver for Power Flow, Optimal Power Flow and State Estimation
Abstract: Power Flow (PF), Optimal Power Flow (OPF), and State Estimation (SE) are fundamental problems in power system analysis, but solving them is computationally expensive. Graph Neural Networks (GNNs) have been proposed as fast surrogates, yet existing solvers are trained for a single problem at a time, producing narrow models that must be rebuilt for each new task. We propose a more general approach: a single Heterogeneous Residual Gated Graph Convolutional Network that solves all three problems with one shared backbone. Rather than learning one mapping, the model learns a reusable representation of how the network behaves, from which PF, OPF, and SE can each be estimated. Trained jointly on the three problems across diverse topologies and loading conditions, and evaluated on the IEEE 14-bus and 118-bus systems, the shared model matches the accuracy of task-specific GNN solvers and stays robust on unseen loading levels and topologies. These results show that a single model can capture the basic operation of a power network and serve several analysis tasks at once, a first step toward a foundation model for power systems.
Physics-aware training improves power flow predictions at scale
Scaling Laws for Physics-Aware ACOPF Surrogate Learning
Abstract: Learning-based surrogates for AC optimal power flow (ACOPF) promise large speedups over classical solvers, but their operational value depends on physical feasibility as much as predictive accuracy. Physics-aware objectives such as the augmented Lagrangian (AL) improve constraint satisfaction at additional per-step cost, yet how this trade-off behaves with scale is uncharacterized. We sweep model and dataset sizes under both MSE and AL training, and characterize how constraint violation changes with network size across grids. Both objectives improve as power laws, but at different rates: MSE is governed primarily by model capacity, while AL is balanced across both. Violation grows roughly twice as fast with network size under MSE as under AL. On matched hardware, AL reduces violation by nearly $30\times$ for an order of magnitude more training time, with negligible added memory. The training objective determines not only where a surrogate lands but how its quality evolves with scale.