Offline reinforcement learning improves wind farm power under changing winds

Offline Reinforcement Learning for Wind Farm Control: A Wind Tunnel Study under Dynamic Wind Directions

Machine Learning

Summary

Wind farms can produce more electricity if turbines face the wind correctly, but wind directions often change. The authors developed a new way to control turbine angles using offline machine learning, which learns from existing data instead of needing lots of trial runs. They tested this method in a wind tunnel experiment and found that it increased power output by about 10% compared to simple controls. This approach also worked as well as state-of-the-art methods that need detailed wind models, while being faster and cheaper to train.

What this means in practice

  • For wind farm operators: Improve overall wind farm power output by adjusting turbine angles using offline-learned control policies that handle changing wind directions.
  • For energy systems engineers: Develop wind turbine control strategies that reduce simulation time and computational cost by using offline data instead of online training.

Authors

Yuhan Su, Hongyang Dong, Simone Tamaro, Filippo Campagnolo, Carlo L. Bottasso, Xiaowei Zhao

Abstract

This paper addresses the wind farm power maximization problem in the presence of wind direction changes. Specifically, a model-free Modified Twin Delayed Deep Deterministic Policy Gradient with Behavior Cloning (MTD3-BC) algorithm is proposed to tackle this task through yaw control under varying wind direction conditions. MTD3-BC is an offline reinforcement learning (RL) algorithm that aims to infer good behavior from only a precollected offline dataset. Additionally, to ensure smooth and moderate yaw adjustments, a new action consistency term is introduced into the policy optimization objective. Unlike online RL methods, MTD3-BC does not require extensive interactions with a wind farm simulator during training, significantly reducing computational costs and training time. A wind tunnel experiment is conducted to validate the effectiveness of the algorithm under varying wind directions. The results demonstrate that MTD3-BC successfully mitigates wake effects, delivering farm-level power gains of approximately 10\% over the baseline greedy strategy and performance on par with a data-calibrated model-based wake-steering benchmark, while requiring no wake model and only a small fraction of the training cost of online RL. To our knowledge, this is the first time an offline RL wind farm control policy has been validated and demonstrated experimentally.