中文

计算马尔可夫决策过程单调策略:一种近保序惩罚方法

系统与控制 2017-04-04 v1

摘要

本文讨论求解具有单调最优策略的马尔可夫决策过程(MDP)的算法。我们提出一种两阶段交替凸优化方案,通过利用单调性质加速最优策略的搜索。第一阶段是以联合状态-动作概率表述的线性规划。第二阶段是以给定状态下动作的条件概率表述的正则化问题。该正则化采用了近保序回归中的技术。尽管在问题的第一种表述中可使用多种迭代方法,但我们在数值模拟中表明,特别是交替乘子法(ADMM)可以通过正则化步骤得到显著加速。

关键词

引用

@article{arxiv.1704.00621,
  title  = {Computing monotone policies for Markov decision processes: a nearly-isotonic penalty approach},
  author = {Robert Mattila and Cristian R. Rojas and Vikram Krishnamurthy and Bo Wahlberg},
  journal= {arXiv preprint arXiv:1704.00621},
  year   = {2017}
}

备注

This work has been accepted for presentation at the 20th World Congress of the International Federation of Automatic Control, 9-14 July 2017