中文
相关论文

相关论文: Hitting time for Markov decision process

200 篇论文

We introduce the avoidance Markov metrics and theories which provide more flexibility in the design of random walk and impose new conditions on the walk to avoid (or transit) a specific node (or a set of nodes) before the stopping criteria.…

离散数学 · 计算机科学 2018-07-18 Golshan Golnari , Zhi-Li Zhang , Daniel Boley

We consider Markov Decision Processes (MDPs) in which every stationary policy induces the same graph structure for the underlying Markov chain and further, the graph has the following property: if we replace each recurrent class by a node,…

机器学习 · 计算机科学 2021-03-10 Joseph Lubars , Anna Winnicki , Michael Livesay , R. Srikant

In this paper, we provide a methodology for computing the probability distribution of sojourn times for a wide class of Markov chains. Our methodology consists in writing out linear systems and matrix equations for generating functions…

概率论 · 数学 2018-01-09 Valentina Cammarota , Aimé Lachal

This paper is devoted to solving a time-inconsistent risk-sensitive control problem with parameter $\e$ and its limit case ($\e\rightarrow0^+$) for countable-stated Markov decision processes (MDPs for short). Since the cost functional is…

最优化与控制 · 数学 2020-10-22 Hongwei Mei

Standard Markov decision process (MDP) and reinforcement learning algorithms optimize the policy with respect to the expected gain. We propose an algorithm which enables to optimize an alternative objective: the probability that the gain is…

机器学习 · 计算机科学 2023-03-06 Vincent Corlay , Jean-Christophe Sibel

Markov decision processes (MDPs) provide a fundamental model for sequential decision making under process uncertainty. A classical synthesis task is to compute for a given MDP a winning policy that achieves a desired specification. However,…

计算机科学中的逻辑 · 计算机科学 2024-07-18 Roman Andriushchenko , Milan Češka , Sebastian Junges , Filip Macák

Markov decision processes (MDPs) are formal models commonly used in sequential decision-making. MDPs capture the stochasticity that may arise, for instance, from imprecise actuators via probabilities in the transition function. However, in…

人工智能 · 计算机科学 2023-06-21 Marnix Suilen , Thiago D. Simão , David Parker , Nils Jansen

Markov automata combine non-determinism, probabilistic branching, and exponentially distributed delays. This compositional variant of continuous-time Markov decision processes is used in reliability engineering, performance evaluation and…

计算机科学中的逻辑 · 计算机科学 2017-05-11 Tim Quatmann , Sebastian Junges , Joost-Pieter Katoen

Quantum walks play an important role in the area of quantum algorithms. Many interesting problems can be reduced to searching marked states in a quantum Markov chain. In this context, the notion of quantum hitting time is very important,…

量子物理 · 物理学 2009-12-08 R. A. M. Santos , R. Portugal

This paper presents a nonparametric method for estimating the conditional density associated to the jump rate of a piecewise-deterministic Markov process. In our framework, the estimation needs only one observation of the process within a…

统计理论 · 数学 2012-07-12 Romain Azaïs , François Dufour , Anne Gégout-Petit

Active classification, i.e., the sequential decision-making process aimed at data acquisition for classification purposes, arises naturally in many applications, including medical diagnosis, intrusion detection, and object tracking. In this…

系统与控制 · 计算机科学 2018-10-02 Bo Wu , Mohamadreza Ahmadi , Suda Bharadwaj , Ufuk Topcu

Stochastic time-varying optimization is an integral part of learning in which the shape of the function changes over time in a non-deterministic manner. This paper considers multiple models of stochastic time variation and analyzes the…

最优化与控制 · 数学 2023-02-23 Ali Yekkehkhany , Han Feng , Donghao Ying , Javad Lavaei

In this paper, we investigate the concentration properties of cumulative reward in Markov Decision Processes (MDPs), focusing on both asymptotic and non-asymptotic settings. We introduce a unified approach to characterize reward…

机器学习 · 计算机科学 2025-12-04 Borna Sayedana , Peter E. Caines , Aditya Mahajan

The convergence, convergence rate and expected hitting time play fundamental roles in the analysis of randomised search heuristics. This paper presents a unified Markov chain approach to studying them. Using the approach, the sufficient and…

最优化与控制 · 数学 2013-12-10 Jun He , Feidun He , Xin Yao

We consider Markov Decision Processes (MDPs) with mean-payoff parity and energy parity objectives. In system design, the parity objective is used to encode \omega-regular specifications, and the mean-payoff and energy objectives can be used…

计算机科学与博弈论 · 计算机科学 2011-04-18 Krishnendu Chatterjee , Laurent Doyen

We revisit the work of Dhar and Majumdar [Phys. Rev. E 59, 6413 (1999)] on the limiting distribution of the temporal mean $M_{t}=t^{-1}\int_{0}^{t}du \sign y_{u}$, for a Gaussian Markovian process $y_{t}$ depending on a parameter $\alpha $,…

统计力学 · 物理学 2016-08-31 G. De Smedt , C. Godreche , J. M. Luck

We consider Markov decision processes (MDP) as generators of sequences of probability distributions over states. A probability distribution is p-synchronizing if the probability mass is at least p in a single state, or in a given set of…

形式语言与自动机理论 · 计算机科学 2018-03-28 Laurent Doyen , Thierry Massart , Mahsa Shirmohammadi

We consider the problem of controlling a Markov decision process (MDP) with a large state space, so as to minimize average cost. Since it is intractable to compete with the optimal policy for large scale problems, we pursue the more modest…

最优化与控制 · 数学 2014-02-28 Yasin Abbasi-Yadkori , Peter L. Bartlett , Alan Malek

Learning a Markov Decision Process (MDP) from a fixed batch of trajectories is a non-trivial task whose outcome's quality depends on both the amount and the diversity of the sampled regions of the state-action space. Yet, many MDPs are…

机器学习 · 计算机科学 2022-03-08 Giorgio Angelotti , Nicolas Drougard , Caroline P. C. Chanel

In this paper we study the distribution of hitting and return times for observations of dynamical systems. We apply this results to get an exponential law for the distribution of hitting and return times for rapidly mixing random dynamical…

动力系统 · 数学 2015-06-19 Jerome Rousseau