马尔可夫决策过程的命中时
机器学习
2022-05-12 v2
摘要
我们定义了马尔可夫决策过程(MDP)的命中时。我们不使用由 MDP 诱导的马尔可夫过程的命中时,因为诱导链可能不具有平稳分布。即使它具有平稳分布,该平稳分布也可能与 MDP 的(归一化)占据测度不一致。我们观察到 MDP 与 PageRank 之间的关系。利用这一观察,我们构造了一个 MP,其平稳分布与 MDP 的归一化占据测度一致,并将 MDP 的命中时定义为该关联 MP 的命中时。
引用
@article{arxiv.2205.03476,
title = {Hitting time for Markov decision process},
author = {Ruichao Jiang and Javad Tavakoli and Yiqinag Zhao},
journal= {arXiv preprint arXiv:2205.03476},
year = {2022}
}
备注
The first version of this paper pointed out some issues in the old version of "Cross-Domain Imitation Learning via Optimal Transport". The authors have then addressed these issues according to our suggestions in a new version. We therefore updated our paper, in which we removed contents related to these issues