English
Related papers

Related papers: Persistent-Transient Policy Evaluation for Markov …

200 papers

We establish a connection between policy evaluation in Markov decision processes and PageRank in network analysis. For a fixed policy, we show that the value function of a discounted Markov decision process can be obtained, up to an…

Optimization and Control · Mathematics 2026-05-04 Konstantin Avrachenkov , Lorenzo Gregoris , Nelly Litvak

In this paper we develop a statistical estimation technique to recover the transition kernel $P$ of a Markov chain $X=(X_m)_{m \in \mathbb N}$ in presence of censored data. We consider the situation where only a sub-sequence of $X$ is…

Statistics Theory · Mathematics 2014-05-05 Flavia Barsotti , Yohann De Castro , Thibault Espinasse , Paul Rochet

A divide-and-conquer approach to analyzing Markov chains (MCs) is not utilized as widely as it could be, despite its potential benefits. One primary reason for this is the fact that most MC decomposition approaches involve a complex and…

Probability · Mathematics 2021-02-25 Katsunobu Sasanuma , Robert Hampshire , Alan Scheller-Wolf

We study non-rectangular robust Markov decision processes under the average-reward criterion, where the ambiguity set couples transition probabilities across states and the adversary commits to a stationary kernel for the entire horizon. We…

Optimization and Control · Mathematics 2026-03-11 Shengbo Wang , Nian Si

Policy gradient methods are extensively used in reinforcement learning as a way to optimize expected return. In this paper, we explore the evolution of the policy parameters, for a special class of exactly solvable POMDPs, as a…

Machine Learning · Computer Science 2020-11-04 Gavin McCracken , Colin Daniels , Rosie Zhao , Anna Brandenberger , Prakash Panangaden , Doina Precup

Given a unichain Markov reward process (MRP), we provide an explicit expression for the bias values in terms of mean first passage times. This result implies a generalization of known Markov chain perturbation bounds for the stationary…

Probability · Mathematics 2024-08-09 Ronald Ortner

In this work, we present the first finite-time analysis of Q-learning with time-varying learning policies (i.e., on-policy sampling) for discounted Markov decision processes under minimal assumptions, requiring only the existence of a…

Machine Learning · Computer Science 2026-04-07 Phalguni Nanda , Zaiwei Chen

Undirected graphical models are widely used in statistics, physics and machine vision. However Bayesian parameter estimation for undirected models is extremely challenging, since evaluation of the posterior typically involves the…

Computation · Statistics 2012-03-19 Richard G. Everitt

Markov chains have been widely employed as a fundamental model in the studies of probabilistic and stochastic communicating and concurrent systems. It is well-understood that decomposition techniques play a key role in reachability analysis…

Quantum Physics · Physics 2018-02-15 Ji Guan , Yuan Feng , Mingsheng Ying

We consider piecewise deterministic Markov processes with degenerate transition kernels of the "house-of-cards"-type. We use a splitting scheme based on jump times to prove the absolute continuity, as well as some regularity, of the…

Probability · Mathematics 2016-01-27 Eva Löcherbach

A Markov decision process can be parameterized by a transition kernel and a reward function. Both play essential roles in the study of reinforcement learning as evidenced by their presence in the Bellman equations. In our inquiry of various…

Machine Learning · Computer Science 2023-09-04 Falcon Z. Dai

Recently regular decision processes have been proposed as a well-behaved form of non-Markov decision process. Regular decision processes are characterised by a transition function and a reward function that depend on the whole history,…

Artificial Intelligence · Computer Science 2022-05-19 Alessandro Ronca , Giuseppe De Giacomo

Traditional reinforcement learning usually assumes either episodic interactions with resets or continuous operation to minimize average or cumulative loss. While episodic settings have many theoretical results, resets are often unrealistic…

Optimization and Control · Mathematics 2026-01-13 Bianca Marin Moreno , Margaux Brégère , Pierre Gaillard , Nadia Oudjane

The standard version of the policy iteration (PI) algorithm fails for semicontinuous models, that is, for models with lower semicontinuous one-step costs and weakly continuous transition law. This is due to the lack of continuity properties…

Optimization and Control · Mathematics 2023-07-17 Óscar Vega-Amaya , Fernando Luque-Vásquez

As is well known, the fundamental matrix $(I - P + e \pi)^{-1}$ plays an important role in the performance analysis of Markov systems, where $P$ is the transition probability matrix, $e$ is the column vector of ones, and $\pi$ is the row…

Optimization and Control · Mathematics 2016-04-26 Li Xia , Peter W. Glynn

In this paper, we investigate a nonparametric approach to provide a recursive estimator of the transition density of a non-stationary piecewise-deterministic Markov process, from only one observation of the path within a long time. In this…

Statistics Theory · Mathematics 2013-05-07 Romain Azaïs

In this paper, we study the controllability and stabilizability properties of the Kolmogorov forward equation of a continuous time Markov chain (CTMC) evolving on a finite state space, using the transition rates as the control parameters.…

Systems and Control · Computer Science 2017-03-29 Karthik Elamvazhuthi , Vaibhav Deshmukh , Matthias Kawski , Spring Berman

We propose a two step strategy for estimating one-dimensional dynamical parameters of a quantum Markov chain, which involves quantum post-processing the output using a coherent quantum absorber and a "pattern counting'' estimator computed…

Quantum Physics · Physics 2025-08-28 Federico Girotti , Alfred Godley , Mădălin Guţă

We study the evaluation of a policy under best- and worst-case perturbations to a Markov decision process (MDP), using transition observations from the original MDP, whether they are generated under the same or a different policy. This is…

Artificial Intelligence · Computer Science 2024-11-05 Andrew Bennett , Nathan Kallus , Miruna Oprescu , Wen Sun , Kaiwen Wang

Consider a sequence (indexed by n) of Markov chains Z^n in R^d characterized by transition kernels that approximately (in n) depend only on the rescaled state n^{-1} Z^n. Subject to a smoothness condition, such a family can be closely…

Probability · Mathematics 2009-08-17 Kamil Szczegot
‹ Prev 1 2 3 10 Next ›