中文

面向电力系统的高计算效率安全强化学习

系统与控制 2022-03-24 v2 系统与控制

摘要

我们提出一种计算高效的安全强化学习(RL)方法,用于高比例可变可再生能源电力系统的频率调节。该方法借鉴集合论控制技术,构建基于神经网络的控制策略,保证满足安全关键状态约束,而无需实时求解模型预测控制或投影问题。通过利用鲁棒受控不变多面体的性质,我们构造了一种新颖的闭式“安全滤波器”,可使用任意基于策略梯度的 RL 算法实现端到端安全学习。随后我们将该安全滤波器与深度确定性策略梯度(DDPG)算法结合,用于修正的 9 总线电力系统频率调节,并表明所学策略比鲁棒线性反馈控制技术更具成本效益,同时保持相同的安全保证。我们还表明,所提范式优于增加了约束违反惩罚的 DDPG。

关键词

引用

@article{arxiv.2110.10333,
  title  = {Computationally Efficient Safe Reinforcement Learning for Power Systems},
  author = {Daniel Tabas and Baosen Zhang},
  journal= {arXiv preprint arXiv:2110.10333},
  year   = {2022}
}

备注

\c{opyright} 2022 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works