中文

具有 Borel 空间与普遍可测策略的马尔可夫决策过程的平均代价最优性不等式

最优化与控制 2020-12-17 v3

摘要

我们考虑具有 Borel 状态与动作空间以及普遍可测策略的平均代价马尔可夫决策过程(MDPs)。对于非负代价模型以及具有 Lyapunov 型稳定性特征的无界代价模型,我们引入一组新条件,在这些条件下通过消失折扣因子方法证明了平均代价最优性不等式(ACOI)。与大多数已有的关于 ACOI 的结果不同,我们的结果不要求 MDPs 的任何紧致性与连续性条件。相反,主要思想是运用 Egoroff 定理所断言的点态收敛可测函数序列的几乎一致收敛性质。我们提出的条件旨在利用这一性质。其中,我们要求对每个状态,在该状态的选定动作子集上,状态转移随机核被有限测度所控制。我们将转移核的这一控制性质与 Egoroff 定理相结合来证明 ACOI。

关键词

引用

@article{arxiv.2001.01357,
  title  = {Average Cost Optimality Inequality for Markov Decision Processes with Borel Spaces and Universally Measurable Policies},
  author = {Huizhen Yu},
  journal= {arXiv preprint arXiv:2001.01357},
  year   = {2020}
}

备注

34 pages; to appear in SIAM Journal on Control and Optimization (this is the accepted version before the galley proof). The contents of this paper consist of (i) the author's results given previously in Section 3 of arXiv:1901.03374v1 and (ii) additional sections for an extended discussion and illustrative examples. arXiv admin note: text overlap with arXiv:1901.03374