English

Persistent-Transient Policy Evaluation for Markov Chains via Minimal Peripheral Quotients

Machine Learning 2026-05-11 v2 Machine Learning Numerical Analysis Numerical Analysis

Abstract

We study fixed-policy evaluation for finite Markov chains that may be reducible and periodic. Classical evaluation methods with gain and bias decomposition are not always diagnostic: the gain records only invariant Ces\`aro averages, while persistent phase-dependent behavior is absorbed into the bias together with genuinely transient effects. We identify the real peripheral invariant subspace K(P)\mathcal{K}(P) of the transition matrix PP as the source of this ambiguity. Quotienting by K(P)\mathcal{K}(P) is the minimal exact quotient that removes all non-decaying modes and makes the remaining dynamics strictly stable. After choosing a gauge projection Π\Pi with kernel K(P)\mathcal{K}(P), the reward admits a unique decomposition r=gΠ+(IP)vΠr = g_\Pi^\star + (I-P)v_\Pi^\star, where gΠg_\Pi^\star is a persistent regime profile and vΠv_\Pi^\star is a gauge-fixed transient component. An exact comparison with classical normalized gain and bias shows that the new pair reallocates the same information so that all persistent modes are represented in gΠg_\Pi^\star and vΠv_\Pi^\star is transient. This decomposition reconstructs finite-horizon returns, recovers statewise average reward, admits a transient-cost interpretation, and yields a stable estimator under a generative model.

Keywords

Cite

@article{arxiv.2602.00474,
  title  = {Persistent-Transient Policy Evaluation for Markov Chains via Minimal Peripheral Quotients},
  author = {Yang Xu and Vaneet Aggarwal},
  journal= {arXiv preprint arXiv:2602.00474},
  year   = {2026}
}
R2 v1 2026-07-01T09:28:59.959Z