A Reinforcement Learning Based Model-Free Wide-Area Damping Control under Random PMU Time Delays
Although wide-area signals sent by remote phasor measurement units (PMUs) provide better solutions to damp inter-area low frequency oscillations, the random time delay of wide-area signal also brings new challenges for the design of wide-area damping controllers (WADCs). Most of existing wide-area damping controllers require either partial or complete knowledge of time delay. Without any knowledge of the random delays of remote PMU signals, this paper presents a reinforcement learning (RL) based model-free WADC to damp the inter-area low frequency oscillations. The RL agent takes the learning process to maximize the reward function by choosing an optimal action on each state. In terms of the RL algorithm, the absolute value of angular velocity difference between two generators is chosen as states while generators' active power set point is selected as actions. Comparison studies are performed on IEEE 10-Generator 39-Bus system with a fixed parameter WADC and a model-based Q-learning WADC. The results show the proposed model-free WADC can quickly damp inter-area low frequency oscillations with random delays of wide-area signals sent by PMUs, while the fixed parameter WADC fails to stable the system and the model-based Q-learning WADC has large steady-state errors.
Paper
Full text
A Reinforcement Learning Based Model-Free Wide-Area Damping Control under Random PMU Time Delays
Semantic Scholar · Engineering · 2021
Abstract
Although wide-area signals sent by remote phasor measurement units (PMUs) provide better solutions to damp inter-area low frequency oscillations, the random time delay of wide-area signal also brings new challenges for the design of wide-area damping controllers (WADCs). Most of existing wide-area damping controllers require either partial or complete knowledge of time delay. Without any knowledge of the random delays of remote PMU signals, this paper presents a reinforcement learning (RL) based model-free WADC to damp the inter-area low frequency oscillations. The RL agent takes the learning process to maximize the reward function by choosing an optimal action on each state. In terms of the RL algorithm, the absolute value of angular velocity difference between two generators is chosen as states while generators' active power set point is selected as actions. Comparison studies are performed on IEEE 10-Generator 39-Bus system with a fixed parameter WADC and a model-based Q-learning WADC. The results show the proposed model-free WADC can quickly damp inter-area low frequency oscillations with random delays of wide-area signals sent by PMUs, while the fixed parameter WADC fails to stable the system and the model-based Q-learning WADC has large steady-state errors.