中文
相关论文

相关论文: A Short Note on Stationary Distributions of Unicha…

200 篇论文

We show that combinations of optimal (stationary) policies in unichain Markov decision processes are optimal. That is, let M be a unichain Markov decision process with state space S, action space A and policies \pi_j^*: S -> A (1\leq j\leq…

组合数学 · 数学 2007-05-23 Ronald Ortner

In this paper, we consider a subclass of piecewise deterministic Markov processes with a Polish state space that involve a deterministic motion punctuated by random jumps, occurring in a Poisson-like fashion with some state-dependent rate,…

概率论 · 数学 2024-05-28 Dawid Czapla

This paper studies Markov Decision Processes (MDPs) with atomless initial state distributions and atomless transition probabilities. Such MDPs are called atomless. The initial state distribution is considered to be fixed. We show that for…

最优化与控制 · 数学 2018-10-26 Eugene A. Feinberg , Aleksey B. Piunovskiy

Constrained Markov Decision Processes (CMDPs) are notably more complex to solve than standard MDPs due to the absence of universally optimal policies across all initial state distributions. This necessitates re-solving the CMDP whenever the…

机器学习 · 计算机科学 2025-10-02 Alperen Tercan , Necmiye Ozay

For a Markov decision process with countably infinite states, the optimal value may not be achievable in the set of stationary policies. In this paper, we study the existence conditions of an optimal stationary policy in a countable-state…

最优化与控制 · 数学 2020-07-06 Li Xia , Xianping Guo , Xi-Ren Cao

We consider average-cost Markov decision processes (MDPs) with Borel state spaces, countable, discrete action spaces, and strictly unbounded one-stage costs. For the minimum pair approach, we introduce a new majorization condition on the…

最优化与控制 · 数学 2020-05-06 Huizhen Yu

We consider the problem of controlling a Markov decision process (MDP) with a large state space, so as to minimize average cost. Since it is intractable to compete with the optimal policy for large scale problems, we pursue the more modest…

最优化与控制 · 数学 2014-02-28 Yasin Abbasi-Yadkori , Peter L. Bartlett , Alan Malek

A classical problem for Markov chains is determining their stationary (or steady-state) distribution. This problem has an equally classical solution based on eigenvectors and linear equation systems. However, this approach does not scale to…

系统与控制 · 电气工程与系统科学 2023-01-20 Tobias Meggendorfer

We consider Markov Decision Processes (MDPs) in which every stationary policy induces the same graph structure for the underlying Markov chain and further, the graph has the following property: if we replace each recurrent class by a node,…

机器学习 · 计算机科学 2021-03-10 Joseph Lubars , Anna Winnicki , Michael Livesay , R. Srikant

The planning domain has experienced increased interest in the formal synthesis of decision-making policies. This formal synthesis typically entails finding a policy which satisfies formal specifications in the form of some well-defined…

人工智能 · 计算机科学 2021-11-30 George K. Atia , Andre Beckus , Ismail Alkhouri , Alvaro Velasquez

Probabilistic model checking can provide formal guarantees on the behavior of stochastic models relating to a wide range of quantitative properties, such as runtime, energy consumption or cost. But decision making is typically with respect…

计算机科学中的逻辑 · 计算机科学 2024-03-19 Ingy Elsayed-Aly , David Parker , Lu Feng

We cast episodic Markov decision process (MDP) planning as Bayesian inference over policies. A policy is treated as the latent variable and is assigned an unnormalized probability of optimality that is monotone in its expected return,…

机器学习 · 计算机科学 2026-04-14 David Tolpin

The parameters of a discrete stationary Markov model are transition probabilities between states. Traditionally, data consist in sequences of observed states for a given number of individuals over the whole observation period. In such a…

统计计算 · 统计学 2012-04-30 Alberto Pasanisi , Shuai Fu , Nicolas Bousquet

We compute the stationary distribution of a continuous-time Markov chain which is constructed by gluing together two finite, irreducible Markov chains by identifying a pair of states of one chain with a pair of states of the other and…

概率论 · 数学 2015-10-22 Bence Mélykúti , Peter Pfaffelhuber

This paper extends to Continuous-Time Jump Markov Decision Processes (CTJMDP) the classic result for Markov Decision Processes stating that, for a given initial state distribution, for every policy there is a (randomized) Markov policy,…

最优化与控制 · 数学 2020-05-18 Eugene A. Feinberg , Manasa Mandava , Albert N. Shiryaev

We consider Markov decision processes (MDP) as generators of sequences of probability distributions over states. A probability distribution is p-synchronizing if the probability mass is at least p in a single state, or in a given set of…

形式语言与自动机理论 · 计算机科学 2018-03-28 Laurent Doyen , Thierry Massart , Mahsa Shirmohammadi

Markov Decision Process (MDP) presents a mathematical framework to formulate the learning processes of agents in reinforcement learning. MDP is limited by the Markovian assumption that a reward only depends on the immediate state and…

机器学习 · 计算机科学 2024-06-04 Bohao Qu , Xiaofeng Cao , Jielong Yang , Hechang Chen , Chang Yi , Ivor W. Tsang , Yew-Soon Ong

Policy gradients in continuous control have been derived for both stochastic and deterministic policies. Here we study the relationship between the two. In a widely-used family of MDPs involving Gaussian control noise and quadratic control…

机器学习 · 计算机科学 2025-06-11 Emo Todorov

The goal of this paper is to analyze distributional Markov Decision Processes as a class of control problems in which the objective is to learn policies that steer the distribution of a cumulative reward toward a prescribed target law,…

最优化与控制 · 数学 2026-02-09 Nicole Bäuerle , Athanasios Vasileiadis

The distributionally robust Markov Decision Process (MDP) approach asks for a distributionally robust policy that achieves the maximal expected total reward under the most adversarial distribution of uncertain parameters. In this paper, we…

系统与控制 · 计算机科学 2018-10-10 Zhi Chen , Pengqian Yu , William B. Haskell
‹ 上一页 1 2 3 10 下一页 ›