English

Comparing discounted and average-cost Markov Decision Processes: a statistical significance perspective

Optimization and Control 2021-12-02 v1 Applications

Abstract

Optimal Markov Decision Process policies for problems with finite state and action space are identified through a partial ordering by comparing the value function across states. This is referred to as state-based optimality. This paper identifies when such optimality guarantees some form of system-based optimality as measured by a scalar. Four such system-based metrics are introduced. Uni-variate empirical distributions of these metrics are obtained through simulation as to assess whether theoretically optimal policies provide a statistically significant advantage. This has been conducted using a Student's t-test, Welch's tt-test and a Mann-Whitney UU-test. The proposed method is applied to a common problem in queuing theory: admission control.

Keywords

Cite

@article{arxiv.2112.00684,
  title  = {Comparing discounted and average-cost Markov Decision Processes: a statistical significance perspective},
  author = {Dylan Solms},
  journal= {arXiv preprint arXiv:2112.00684},
  year   = {2021}
}
R2 v1 2026-06-24T08:00:06.994Z