Comprehensive benchmark improves wind power forecasting evaluation
WPBench: A Comprehensive Benchmark for Wind Power Forecasting
Machine LearningArtificial Intelligence
Summary
Predicting how much wind power will be produced is important for managing electricity and integrating renewable energy. Current ways to test these predictions have gaps, such as not covering all types of wind turbines or models. The authors created WPBench, a big testing system that uses many datasets and models, evaluates multiple aspects of predictions, and helps compare methods fairly. This tool makes it easier to see which forecasting methods work best across many wind scenarios.
What this means in practice
- •For power grid operators: Assess and select wind power forecasting models using unified benchmarks that cover various turbine scales and conditions for more reliable grid management.
- •For renewable energy software developers: Develop and test forecasting software with WPBench to ensure compatibility with diverse datasets and evaluation criteria, improving deployment readiness.
Authors
Yuhan Zhu, Jilin Hu, Xinying Cai, Yingshan Li, Li Ma, Xiangfei Qiu Linsen Li, Kai Zhang, Yao Fu, Weihao Jiang, Bin Yang
Abstract
Accurate, reliable, and deployable wind power forecasting is critical for power system dispatch, renewable energy integration, and electricity market operations. Progress in this field hinges on the ability to empirically and comprehensively benchmark forecasting methods. Yet existing benchmarks fall short of supporting systematic evaluation in four key aspects: 1) limited coverage of wind power scenarios across turbine scale, variable composition, and spatial structure; 2) incomplete coverage of forecasting model families; 3) evaluation metrics misaligned with wind power requirements; and 4) limited structure-aware diagnostics beyond individual temporal patterns. To address these limitations, we propose WPBench, a comprehensive, fair, and extensible benchmark for wind power forecasting. WPBench integrates 26 public datasets organized by turbine scale and variable composition, spanning single-turbine, multi-turbine, univariate, and multivariate settings. Under unified processing, training, and evaluation protocols, it benchmarks 19 representative models covering traditional methods, deep temporal models, spatio-temporal models, and foundation models. Beyond point-wise errors, WPBench assesses forecast-curve fidelity and computational efficiency, and delivers structure-aware diagnostics across temporal, variable-dependency, and spatial-dependency perspectives. Together, these capabilities enable systematic model comparison across diverse wind scenarios and provide a reusable platform for future research.