计算马尔可夫决策过程单调策略:一种近保序惩罚方法
系统与控制
2017-04-04 v1
摘要
本文讨论求解具有单调最优策略的马尔可夫决策过程(MDP)的算法。我们提出一种两阶段交替凸优化方案,通过利用单调性质加速最优策略的搜索。第一阶段是以联合状态-动作概率表述的线性规划。第二阶段是以给定状态下动作的条件概率表述的正则化问题。该正则化采用了近保序回归中的技术。尽管在问题的第一种表述中可使用多种迭代方法,但我们在数值模拟中表明,特别是交替乘子法(ADMM)可以通过正则化步骤得到显著加速。
引用
@article{arxiv.1704.00621,
title = {Computing monotone policies for Markov decision processes: a nearly-isotonic penalty approach},
author = {Robert Mattila and Cristian R. Rojas and Vikram Krishnamurthy and Bo Wahlberg},
journal= {arXiv preprint arXiv:1704.00621},
year = {2017}
}
备注
This work has been accepted for presentation at the 20th World Congress of the International Federation of Automatic Control, 9-14 July 2017