中文

具有指数效用连续时间马尔可夫决策过程的渐进-脉冲控制

最优化与控制 2023-11-16 v2

摘要

在本文中,我们考虑连续时间马尔可夫决策过程的渐进-脉冲控制问题,其中系统性能由总成本的指数效用的期望来度量。在系统基本要素非常一般的条件下,我们证明了在一类更一般的策略中存在确定性平稳最优策略。我们所考虑的策略允许多个同时脉冲、具有随机效应的脉冲的随机化选择、松弛渐进控制以及跳跃的累积。在利用最优方程刻画值函数之后,我们将连续时间渐进-脉冲控制问题化为一个等价的简单离散时间马尔可夫决策过程,其动作空间是渐进动作集与脉冲动作集的并集。

关键词

引用

@article{arxiv.1811.11704,
  title  = {On gradual-impulse control of continuous-time Markov decision processes with exponential utility},
  author = {Xin Guo and Aiko Kurushima and Alexey Piunovskiy and Yi Zhang},
  journal= {arXiv preprint arXiv:1811.11704},
  year   = {2023}
}

备注

The proof of Lemma 4 in the published version was wrong (though the statement is correct). It is now corrected it (see the proof of Lemma 5.4 in the current file)