Strategies in POMDPs with Stage Duration
Abstract
Partially observable Markov decision processes (POMDPs) with stage duration provide a framework for approximating continuous-time behavior by scaling transition probabilities with a stage duration parameter . While previous literature has primarily focused on the limit of the discounted value as the stage duration vanishes, this paper investigates the global behavior of the asymptotic value, , across varying stage durations. Our main result demonstrates that any strategy in a POMDP with stage duration can be mimicked in the base POMDP (). Specifically, we provide an explicit construction showing that for any strategy in the POMDP with stage duration , there exists a strategy in the base POMDP that secures the same asymptotic payoff. As a consequence of this theorem, we establish that the value function is nondecreasing with respect to , and that the continuous-time limit exists.
Cite
@article{arxiv.2603.16055,
title = {Strategies in POMDPs with Stage Duration},
author = {Ivan Novikov},
journal= {arXiv preprint arXiv:2603.16055},
year = {2026}
}