通过搜索有限策略空间求解部分可观测马尔可夫决策过程
人工智能
2013-01-30 v1
摘要
求解部分可观测马尔可夫决策过程(POMDP)在一般情况下极为困难,部分原因在于最优策略可能具有无限规模。本文探讨了从受限策略集合中寻找最优策略的问题,该集合表示为给定规模的有限状态自动机。该问题同样难以求解,但我们证明,当对POMDP和/或策略施加进一步约束时,复杂度可大幅降低。我们展示了使用分支定界法寻找全局最优确定性策略,以及使用梯度上升法寻找局部最优随机策略的良好实证结果。
引用
@article{arxiv.1301.6720,
title = {Solving POMDPs by Searching the Space of Finite Policies},
author = {Nicolas Meuleau and Kee-Eung Kim and Leslie Pack Kaelbling and Anthony R. Cassandra},
journal= {arXiv preprint arXiv:1301.6720},
year = {2013}
}
备注
Appears in Proceedings of the Fifteenth Conference on Uncertainty in Artificial Intelligence (UAI1999)