English

Geometry of Drifting MDPs with Path-Integral Stability Certificates

Machine Learning 2026-01-30 v1 Artificial Intelligence

Abstract

Real-world reinforcement learning is often \emph{nonstationary}: rewards and dynamics drift, accelerate, oscillate, and trigger abrupt switches in the optimal action. Existing theory often represents nonstationarity with coarse-scale models that measure \emph{how much} the environment changes, not \emph{how} it changes locally -- even though acceleration and near-ties drive tracking error and policy chattering. We take a geometric view of nonstationary discounted Markov Decision Processes (MDPs) by modeling the environment as a differentiable homotopy path and tracking the induced motion of the optimal Bellman fixed point. This yields a length-curvature-kink signature of intrinsic complexity: cumulative drift, acceleration/oscillation, and action-gap-induced nonsmoothness. We prove a solver-agnostic path-integral stability bound and derive gap-safe feasible regions that certify local stability away from switch regimes. Building on these results, we introduce \textit{Homotopy-Tracking RL (HT-RL)} and \textit{HT-MCTS}, lightweight wrappers that estimate replay-based proxies of length, curvature, and near-tie proximity online and adapt learning or planning intensity accordingly. Experiments show improved tracking and dynamic regret over matched static baselines, with the largest gains in oscillatory and switch-prone regimes.

Keywords

Cite

@article{arxiv.2601.21991,
  title  = {Geometry of Drifting MDPs with Path-Integral Stability Certificates},
  author = {Zuyuan Zhang and Mahdi Imani and Tian Lan},
  journal= {arXiv preprint arXiv:2601.21991},
  year   = {2026}
}
R2 v1 2026-07-01T09:26:08.478Z