English
Related papers

Related papers: Persistent-Transient Policy Evaluation for Markov …

200 papers

We introduce a unified operator-theoretic framework for analyzing mixing times of finite-state ergodic Markov chains that applies to both reversible and non-reversible dynamics. The central object in our analysis is the projected transition…

Probability · Mathematics 2025-11-05 Muhammad Abdullah Naeem

Markov chain analysis is a key technique in formal verification. A practical obstacle is that all probabilities in Markov models need to be known. However, system quantities such as failure rates or packet loss ratios, etc. are often not --…

Logic in Computer Science · Computer Science 2023-11-08 Sebastian Junges , Erika Ábrahám , Christian Hensel , Nils Jansen , Joost-Pieter Katoen , Tim Quatmann , Matthias Volk

We consider nonparametric estimation of the transition operator $P$ of a Markov chain and its transition density $p$ where the singular values of $P$ are assumed to decay exponentially fast. This is for instance the case for periodised,…

Statistics Theory · Mathematics 2021-10-26 Matthias Löffler , Antoine Picard

Experience replay is a core ingredient of modern deep reinforcement learning, yet its benefits in policy optimization are poorly understood beyond empirical heuristics. This paper develops a novel theoretical framework for experience replay…

Machine Learning · Computer Science 2026-02-04 Hua Zheng , Wei Xie , M. Ben Feng

Policy gradient methods in reinforcement learning update policy parameters by taking steps in the direction of an estimated gradient of policy value. In this paper, we consider the statistically efficient estimation of policy gradients from…

Machine Learning · Statistics 2020-02-21 Nathan Kallus , Masatoshi Uehara

A wide variety of physical systems ranging from the firing of neurons to eutrophication of lakes to the presence of Arctic summer sea ice exhibit a phenomenon known as tipping. In mathematical models, tipping can be caused by bifurcations,…

Dynamical Systems · Mathematics 2018-03-14 Alanna Hoyer-Leitzel , Alice Nadeau , Andrew Roberts , Andrew Steyer

This paper is concerned with the development of rigorous approximations to various expectations associated with Markov chains and processes having non-stationary transition probabilities. Such non-stationary models arise naturally in…

Probability · Mathematics 2018-05-07 Zeyu Zheng , Harsha Honnappa , Peter W. Glynn

In this paper, we study a mean-variance optimization problem in an infinite horizon discrete time discounted Markov decision process (MDP). The objective is to minimize the variance of system rewards with the constraint of mean performance.…

Optimization and Control · Mathematics 2017-08-24 Li Xia

We provide a novel method for sensitivity analysis of parametric robust Markov chains. These models incorporate parameters and sets of probability distributions to alleviate the often unrealistic assumption that precise probabilities are…

Machine Learning · Computer Science 2023-05-03 Thom Badings , Sebastian Junges , Ahmadreza Marandi , Ufuk Topcu , Nils Jansen

We study random walks on contingency tables with fixed marginals, corresponding to a (log-linear) hierarchical model. If the set of allowed moves is not a Markov basis, then there exist tables with the same marginals that are not connected.…

Commutative Algebra · Mathematics 2016-04-08 Thomas Kahle , Johannes Rauh , Seth Sullivant

Accurate characterization of coherent and non-Markovian errors remains a central challenge in quantum information processing, as conventional benchmarking techniques typically rely on Markovian and time-independent noise assumptions. In…

We propose a new approach for estimating the finite dimensional transition matrix of a Markov chain using a large number of independent sample paths observed at random times. The sample paths may be observed as few as two times, and the…

Methodology · Statistics 2025-05-20 Daphne Aurouet , Valentin Patilea

Markov-modulated Brownian motion is a popular tool to model continuous-time phenomena in a stochastic context. The main quantity of interest is the invariant density, which satisfies a differential equation associated with the quadratic…

Probability · Mathematics 2016-05-06 Giang T. Nguyen , Federico Poloni

The spectral gap of a Markov chain can be bounded by the spectral gaps of constituent "restriction" chains and a "projection" chain, and the strength of such a bound is the content of various decomposition theorems. In this paper, we…

Data Structures and Algorithms · Computer Science 2019-10-14 Sarah Miracle , Amanda Pascoe Streib , Noah Streib

Motivated by applications arising in networked systems, this work examines controlled regime-switching systems that stem from a mean-variance formulation. A main point is that the switching process is a hidden Markov chain. An additional…

Optimization and Control · Mathematics 2014-01-21 Zhixin Yang , George Yin , Qing Zhang

We present a coherent approach to recurrence and transience, starting from a version of the Riesz decomposition theorem for superharmonic elements. Our approach allows straightforward proofs of some known results, entails new theorems, and…

Operator Algebras · Mathematics 2012-11-30 Andreas Gärtner , Burkhard Kümmerer

We consider Markov Decision Processes (MDPs) in which every stationary policy induces the same graph structure for the underlying Markov chain and further, the graph has the following property: if we replace each recurrent class by a node,…

Machine Learning · Computer Science 2021-03-10 Joseph Lubars , Anna Winnicki , Michael Livesay , R. Srikant

We study finite horizon optimal switching problems for hidden Markov chain models under partially observable Poisson processes. The controller possesses a finite range of strategies and attempts to track the state of the unobserved state…

Optimization and Control · Mathematics 2008-05-22 Erhan Bayraktar , Mike Ludkovski

Consider a sequence $(\eta^N(t) :t\ge 0)$ of continuous-time, irreducible Markov chains evolving on a fixed finite set $E$, indexed by a parameter $N$. Denote by $R_N(\eta,\xi)$ the jump rates of the Markov chain $\eta^N_t$, and assume that…

Probability · Mathematics 2015-12-22 C. Landim , T. Xu

This paper investigates recursive feasibility, recursive robust stability and near-optimality properties of policy iteration (PI). For this purpose, we consider deterministic nonlinear discrete-time systems whose inputs are generated by PI…

Optimization and Control · Mathematics 2022-10-27 Mathieu Granzotto , Olivier Lindamulage De Silva , Romain Postoyan , Dragan Nesic , Zhong-Ping Jiang