面向 POMDP 的强化学习:分区滚动策略与策略迭代及其在自主顺序维修问题中的应用
机器人学
2020-02-12 v1 人工智能
摘要
在本文中,我们考虑具有有限状态与控制空间以及部分状态观测的无限 horizon 折扣动态规划问题。我们讨论了一种采用多步前瞻、基于已知基策略的截断滚动(rollout)以及终端代价函数近似的算法。该算法亦用于近似策略迭代框架中的策略改进,其中相继策略通过神经网络分类器近似。我们方法的新颖之处在于,它通过扩展信念空间表述与分区架构(由多个神经网络训练)非常适于分布式计算。我们在仿真中将方法应用于一类顺序维修问题:机器人在关于管道状态的部分信息下检查并修复可能具有多个破裂点的管道。
引用
@article{arxiv.2002.04175,
title = {Reinforcement Learning for POMDP: Partitioned Rollout and Policy Iteration with Application to Autonomous Sequential Repair Problems},
author = {Sushmita Bhattacharya and Sahil Badyal and Thomas Wheeler and Stephanie Gil and Dimitri Bertsekas},
journal= {arXiv preprint arXiv:2002.04175},
year = {2020}
}
备注
Total 9 pages, 9 figures, 1 table, submitted and accepted to be published in IEEE RA-L 2020