关于具有Lipschitz有界策略网络的鲁棒强化学习
机器学习
2025-02-07 v3 系统与控制
系统与控制
摘要
本文对深度强化学习中的鲁棒策略网络进行了研究。我们探究了自然满足其Lipschitz界约束的策略参数化的优势,并在两个代表性问题上分析了其实验性能和鲁棒性:摆杆起摆和Atari Pong。我们说明,与由普通多层感知机或卷积神经网络组成的无约束策略相比,具有较小Lipschitz界的策略网络对干扰、随机噪声和针对性对抗攻击更为鲁棒。然而,Lipschitz层的结构很重要。我们发现广泛使用的谱归一化方法过于保守,严重影响了干净性能,而更具表达力的Lipschitz层(如最近提出的Sandwich层)可以在不牺牲干净性能的情况下实现更好的鲁棒性。
引用
@article{arxiv.2405.11432,
title = {On Robust Reinforcement Learning with Lipschitz-Bounded Policy Networks},
author = {Nicholas H. Barbara and Ruigang Wang and Ian R. Manchester},
journal= {arXiv preprint arXiv:2405.11432},
year = {2025}
}
备注
Accepted to the Symposium on Systems Theory in Data and Optimization (SysDO 2024)