中文
相关论文

相关论文: Hitting time for Markov decision process

200 篇论文

Consider a discrete time, ergodic Markov chain with finite state space which is started from stationarity. Fill and Lyzinski (2014) showed that, in some cases, the hitting time for a given state may be represented as a sum of a geometric…

概率论 · 数学 2018-12-20 Fraser Daly

In the absence of acceleration, the velocity formula gives "distance travelled equals speed multiplied by time". For a broad class of Markov chains such as circulant Markov chains or random walk on complete graphs, we prove a probabilistic…

概率论 · 数学 2018-06-20 Michael C. H. Choi

We develop some sufficient conditions for the stochastic ordering between hitting times, in a fixed state, for two Markov chains. In particular, we focus attention on the so called \emph{skip-free} case. In the analysis of such a case, we…

概率论 · 数学 2014-03-25 Emilio De Santis , Fabio Spizzichino

This study introduces a comparative modeling framework using stationary and non-stationary transition probabilities within a Markov Decision Process (MDP) to assess COVID-19 disease dynamics. Stationary transition probabilities assume…

Markov decision problems (MDPs) provide the foundations for a number of problems of interest to AI researchers studying automated planning and reinforcement learning. In this paper, we summarize results regarding the complexity of solving…

人工智能 · 计算机科学 2013-02-21 Michael L. Littman , Thomas L. Dean , Leslie Pack Kaelbling

This paper considers an infinite-horizon Markov decision process (MDP) that allows for general non-exponential discount functions, in both discrete and continuous time. Due to the inherent time inconsistency, we look for a randomized…

最优化与控制 · 数学 2024-12-10 Erhan Bayraktar , Yu-Jui Huang , Zhenhua Wang , Zhou Zhou

We investigate the hitting times of random walks on graphs, where a hitting time is defined as the number of steps required for a random walker to move from one node to another. While much of the existing literature focuses on calculating…

概率论 · 数学 2025-11-10 Anuraag Kumar

We establish a connection between policy evaluation in Markov decision processes and PageRank in network analysis. For a fixed policy, we show that the value function of a discounted Markov decision process can be obtained, up to an…

最优化与控制 · 数学 2026-05-04 Konstantin Avrachenkov , Lorenzo Gregoris , Nelly Litvak

This paper studies parametric Markov decision processes (pMDPs), an extension to Markov decision processes (MDPs) where transitions probabilities are described by polynomials over a finite set of parameters. Fixing values for all parameters…

计算机科学中的逻辑 · 计算机科学 2019-04-03 Tobias Winkler , Sebastian Junges , Guillermo A. Pérez , Joost-Pieter Katoen

We consider reinforcement learning for continuous-time Markov decision processes (MDPs) in the infinite-horizon, average-reward setting. In contrast to discrete-time MDPs, a continuous-time process moves to a state and stays there for a…

机器学习 · 计算机科学 2024-07-03 Xuefeng Gao , Xun Yu Zhou

In the setting of non-reversible Markov chains on finite or countable state space, exact results on the distribution of the first hitting time to a given set $G$ are obtained. A new notion of "strong metastability time" is introduced to…

概率论 · 数学 2018-08-01 F. Manzo , E. Scoppola

In the classical theory of Markov chains, one may study the mean time to reach some chosen state, and it is well-known that in the irreducible, finite case, such quantity can be calculated in terms of the fundamental matrix of the walk, as…

量子物理 · 物理学 2022-06-17 C. F. Lardizabal , L. Velázquez

Markov decision processes (MDPs) are standard models for probabilistic systems with non-deterministic behaviours. Mean payoff (or long-run average reward) provides a mathematically elegant formalism to express performance related…

性能 · 计算机科学 2017-09-08 Jan Křetínský , Tobias Meggendorfer

A labelled Markov decision process is a labelled Markov chain with nondeterminism, i.e., together with a strategy a labelled MDP induces a labelled Markov chain. The model is related to interval Markov chains. Motivated by applications of…

形式语言与自动机理论 · 计算机科学 2020-09-25 Stefan Kiefer , Qiyi Tang

We consider the problem of characterising expected hitting times and hitting probabilities for imprecise Markov chains. To this end, we consider three distinct ways in which imprecise Markov chains have been defined in the literature: as…

概率论 · 数学 2020-01-28 Thomas Krak , Natan T'Joens , Jasper De Bock

Given an irreducible discrete-time Markov chain on a finite state space, we consider the largest expected hitting time $T(\alpha)$ of a set of stationary measure at least $\alpha$ for $\alpha\in(0,1)$. We obtain tight inequalities among the…

The distribution of the "mixing time" or the "time to stationarity" in a discrete time irreducible Markov chain, starting in state i, can be defined as the number of trials to reach a state sampled from the stationary distribution of the…

概率论 · 数学 2014-03-05 Jeffrey J. Hunter

This paper is dedicated to the numerical study of the optimization of an industrial launcher integration process. It is an original case of inventory-production system where a calendar plays a crucial role. The process is modeled using the…

This paper investigates the optimization problem of an infinite stage discrete time Markov decision process (MDP) with a long-run average metric considering both mean and variance of rewards together. Such performance metric is important…

最优化与控制 · 数学 2020-08-11 Li Xia

In this paper, we consider the optimal stopping problem on semi-Markov processes (SMPs) with finite horizon, and aim to establish the existence and computation of optimal stopping times. To achieve the goal, we first develop the main…

概率论 · 数学 2021-07-16 Fang Chen , Xianping Guo , Zhong-Wei Liao