中文
相关论文

相关论文: Hitting time for Markov decision process

200 篇论文

Markov decision processes (MDP) are a well-established model for sequential decision-making in the presence of probabilities. In robust MDP (RMDP), every action is associated with an uncertainty set of probability distributions, modelling…

人工智能 · 计算机科学 2024-12-16 Tobias Meggendorfer , Maximilian Weininger , Patrick Wienhöft

A well-known theorem for an irreducible skip-free chain with absorbing state $d$, under some conditions, is that the hitting (absorbing) time of state $d$ starting from state 0 is distributed as the sum of $d$ independent geometric (or…

概率论 · 数学 2013-01-31 Wenming Hong , Ke Zhou

For any discrete target distribution, we exploit the connection between Markov chains and Stein's method via the generator approach and express the solution of Stein's equation in terms of expected hitting time. This yields new upper bounds…

概率论 · 数学 2018-02-16 Michael C. H. Choi

We introduce synchronizing objectives for Markov decision processes (MDP). Intuitively, a synchronizing objective requires that eventually, at every step there is a state which concentrates almost all the probability mass. In particular, it…

计算机科学中的逻辑 · 计算机科学 2011-02-22 Laurent Doyen , Thierry Massart , Mahsa Shirmohammadi

The hitting time is the required minimum time for a Markov chain-based walk (classical or quantum) to reach a target state in the state space. We investigate the effect of the perturbation on the hitting time of a quantum walk. We obtain an…

量子物理 · 物理学 2013-06-12 Chen-Fu Chiang , Guillermo Gomez

Given a discrete source distribution $\mu$ and discrete target distribution $\nu$ on a common finite state space $\mathcal{X}$, we are tasked with transporting $\mu$ to $\nu$ using a given discrete-time Markov chain $X$ with the quickest…

概率论 · 数学 2018-07-23 Michael C. H. Choi

This paper attempts to study the optimal stopping time for semi-Markov processes (SMPs) under the discount optimization criteria with unbounded cost rates. In our work, we introduce an explicit construction of the equivalent semi-Markov…

概率论 · 数学 2021-01-05 Fang Chen , Xianping Guo , Zhong-Wei Liao

Let 0<\alpha<1/2. We show that the mixing time of a continuous-time reversible Markov chain on a finite state space is about as large as the largest expected hitting time of a subset of stationary measure at least \alpha of the state space.…

概率论 · 数学 2012-08-28 Roberto Imbuzeiro Oliveira

Markov Decision Process (MDP) presents a mathematical framework to formulate the learning processes of agents in reinforcement learning. MDP is limited by the Markovian assumption that a reward only depends on the immediate state and…

机器学习 · 计算机科学 2024-06-04 Bohao Qu , Xiaofeng Cao , Jielong Yang , Hechang Chen , Chang Yi , Ivor W. Tsang , Yew-Soon Ong

Structural results impose sufficient conditions on the model parameters of a Markov decision process (MDP) so that the optimal policy is an increasing function of the underlying state. The classical assumptions for MDP structural results…

系统与控制 · 电气工程与系统科学 2023-03-07 Vikram Krishnamurthy

We introduce the spatiotemporal Markov decision process (STMDP), a special type of Markov decision process that models sequential decision-making problems which are not only characterized by temporal, but also by spatial interaction…

最优化与控制 · 数学 2025-01-08 M. C. de Jongh , Richard J. Boucherie , M. N. M. van Lieshout

Consider a sequence of continuous-time irreducible reversible Markov chains and a sequence of initial distributions, $\mu_n$. The sequence is said to exhibit $\mu_n$-cutoff if the convergence to stationarity in total variation distance is…

概率论 · 数学 2018-02-27 Jonathan Hermon

Markov Decision Processes (MDPs) have been used to formulate many decision-making problems in science and engineering. The objective is to synthesize the best decision (action selection) policies to maximize expected rewards (or minimize…

最优化与控制 · 数学 2015-07-07 Mahmoud El Chamie , Behcet Acikmese

The Markov decision process (MDP) formulation used to model many real-world sequential decision making problems does not efficiently capture the setting where the set of available decisions (actions) at each time step is stochastic.…

机器学习 · 计算机科学 2020-01-22 Yash Chandak , Georgios Theocharous , Blossom Metevier , Philip S. Thomas

We present a general framework for applying learning algorithms and heuristical guidance to the verification of Markov decision processes (MDPs). The primary goal of our techniques is to improve performance by avoiding an exhaustive…

This study introduces a novel approach for learning mixtures of Markov chains, a critical process applicable to various fields, including healthcare and the analysis of web users. Existing research has identified a clear divide in…

机器学习 · 计算机科学 2024-05-27 Fabian Spaeh , Konstantinos Sotiropoulos , Charalampos E. Tsourakakis

Consider a Markov chain with finite state space and suppose you wish to change time replacing the integer step index $n$ with a random counting process $N(t)$. What happens to the mixing time of the Markov chain? We present a partial reply…

概率论 · 数学 2021-11-17 Nicos Georgiou , Enrico Scalas

We study the hitting times of Markov processes to target set $G$, starting from a reference configuration $x_0$ or its basin of attraction. The configuration $x_0$ can correspond to the bottom of a (meta)stable well, while the target $G$…

概率论 · 数学 2014-06-11 R. Fernandez , F. Manzo , F. R. Nardi , E. Scoppola

The goal of a traditional Markov decision process (MDP) is to maximize expected cumulative reward over a defined horizon (possibly infinite). In many applications, however, a decision maker may be interested in optimizing a specific…

人工智能 · 计算机科学 2025-10-16 Xiaocheng Li , Huaiyang Zhong , Margaret L. Brandeau

Markov decision processes (MDPs) are used to model a wide variety of applications ranging from game playing over robotics to finance. Their optimal policy typically maximizes the expected sum of rewards given at each step of the decision…

机器学习 · 计算机科学 2025-05-26 Maximilian Nägele , Jan Olle , Thomas Fösel , Remmy Zen , Florian Marquardt