带乘性噪声离散时间系统随机 LQ 控制的强化学习
最优化与控制
2023-11-22 v1
摘要
本文考虑无限时域下带乘性噪声的离散时间系统的随机线性二次问题。为获得最优解,我们提出了一种基于贝尔曼动态规划原理的在线迭代强化学习算法。该算法避免直接求解代数黎卡提方程(algebraic Riccati equations)。它仅利用短区间上的状态轨迹而非所有迭代,显著简化了计算过程。在可镇定的初值下,数值算例阐明了我们的理论结果。
引用
@article{arxiv.2311.12322,
title = {Reinforcement Learning for Stochastic LQ Control of Discrete-Time Systems with Multiplicative Noises},
author = {Hongdan Li and Lucky Qiaofeng Li and Xun Li and Zhaorong Zhang},
journal= {arXiv preprint arXiv:2311.12322},
year = {2023}
}