中文

通过搜索有限策略空间求解部分可观测马尔可夫决策过程

人工智能 2013-01-30 v1

摘要

求解部分可观测马尔可夫决策过程(POMDP)在一般情况下极为困难,部分原因在于最优策略可能具有无限规模。本文探讨了从受限策略集合中寻找最优策略的问题,该集合表示为给定规模的有限状态自动机。该问题同样难以求解,但我们证明,当对POMDP和/或策略施加进一步约束时,复杂度可大幅降低。我们展示了使用分支定界法寻找全局最优确定性策略,以及使用梯度上升法寻找局部最优随机策略的良好实证结果。

关键词

引用

@article{arxiv.1301.6720,
  title  = {Solving POMDPs by Searching the Space of Finite Policies},
  author = {Nicolas Meuleau and Kee-Eung Kim and Leslie Pack Kaelbling and Anthony R. Cassandra},
  journal= {arXiv preprint arXiv:1301.6720},
  year   = {2013}
}

备注

Appears in Proceedings of the Fifteenth Conference on Uncertainty in Artificial Intelligence (UAI1999)