有限时域连续强化学习的一种易处理算法
机器学习
2019-08-05 v1 人工智能
摘要
我们考虑有限时域连续强化学习问题。本文贡献有三方面。首先,我们针对该问题给出一种基于乐观值迭代的易处理算法。其次,我们给出对于任何离散化状态空间的算法,其遗憾下界为 量级,改进了 Ortner 和 Ryabko \cite{contrl} 针对同一问题先前给出的 遗憾界。接着,在奖励和转移满足 Hölder 连续的假设下,我们证明离散化误差的上界为 。最后,我们给出一些简单实验以验证我们的命题。
引用
@article{arxiv.1906.11245,
title = {A Tractable Algorithm For Finite-Horizon Continuous Reinforcement Learning},
author = {Phanideep Gampa and Sairam Satwik Kondamudi and Lakshmanan Kailasam},
journal= {arXiv preprint arXiv:1906.11245},
year = {2019}
}
备注
InProceedings of International Conference on Intelligent Autonomous System, ICOIAS 2019