English

Beyond the Bellman Fixed Point: Geometry and Fast Policy Identification in Value Iteration

Optimization and Control 2026-05-06 v4 Artificial Intelligence Systems and Control Systems and Control

Abstract

Q-value iteration (Q-VI) is usually analyzed through the γ\gamma-contraction of the Bellman operator. This argument proves convergence to QQ^*, but it gives only a coarse account of when the induced greedy policy becomes optimal. We study discounted Q-VI as a switching system and focus on the practically optimal solution set (POSS), the set of QQ-functions whose tie-broken greedy policies are optimal. The main result shows that Q-VI reaches the optimal action class in finite time by entering an invariant tube around X1=Q+span(1)\mathcal X_1=Q^*+\operatorname{span}(\mathbf 1), which is contained in the POSS. For every ε>0\varepsilon>0, the distance to X1\mathcal X_1 satisfies an exponential bound with rate (ρˉ+ε)k(\bar\rho+\varepsilon)^k, where ρˉ\bar\rho is the joint spectral radius of the projected switching family restricted to directions transverse to X1\mathcal X_1. When ρˉ<γ\bar\rho<\gamma, this transverse convergence is faster than the classical contraction rate. The analysis separates fast policy identification from the subsequent convergence to QQ^*, which may still be governed by the all-ones mode. We also give spectral and graph-theoretic conditions under which the strict inequality ρˉ<γ\bar\rho<\gamma holds or fails.

Keywords

Cite

@article{arxiv.2604.17457,
  title  = {Beyond the Bellman Fixed Point: Geometry and Fast Policy Identification in Value Iteration},
  author = {Donghwan Lee},
  journal= {arXiv preprint arXiv:2604.17457},
  year   = {2026}
}
R2 v1 2026-07-01T12:16:56.899Z