English

A Generalized Fundamental Matrix for Computing Fundamental Quantities of Markov Systems

Optimization and Control 2016-04-26 v2 Systems and Control

Abstract

As is well known, the fundamental matrix (IP+eπ)1(I - P + e \pi)^{-1} plays an important role in the performance analysis of Markov systems, where PP is the transition probability matrix, ee is the column vector of ones, and π\pi is the row vector of the steady state distribution. It is used to compute the performance potential (relative value function) of Markov decision processes under the average criterion, such as g=(IP+eπ)1fg=(I - P + e \pi)^{-1} f where gg is the column vector of performance potentials and ff is the column vector of reward functions. However, we need to pre-compute π\pi before we can compute (IP+eπ)1(I - P + e \pi)^{-1}. In this paper, we derive a generalization version of the fundamental matrix as (IP+er)1(I - P + e r)^{-1}, where rr can be any given row vector satisfying re0r e \neq 0. With this generalized fundamental matrix, we can compute g=(IP+er)1fg=(I - P + e r)^{-1} f. The steady state distribution is computed as π=r(IP+er)1\pi = r(I - P + e r)^{-1}. The Q-factors at every state-action pair can also be computed in a similar way. These formulas may give some insights on further understanding how to efficiently compute or estimate the values of gg, π\pi, and Q-factors in Markov systems, which are fundamental quantities for the performance optimization of Markov systems.

Cite

@article{arxiv.1604.04343,
  title  = {A Generalized Fundamental Matrix for Computing Fundamental Quantities of Markov Systems},
  author = {Li Xia and Peter W. Glynn},
  journal= {arXiv preprint arXiv:1604.04343},
  year   = {2016}
}
R2 v1 2026-06-22T13:32:57.841Z