中文
相关论文

相关论文: Hitting time for Markov decision process

200 篇论文

The hitting and mixing times are two fundamental quantities associated with Markov chains. In Peres and Sousi[PS2015] and Oliveira[Oli2012], the authors show that the mixing times and "worst-case" hitting times of reversible Markov chains…

概率论 · 数学 2019-04-05 Robert M. Anderson , Haosui Duanmu , Aaron Smith

In spiking neural networks, the information is conveyed by the spike times, that depend on the intrinsic dynamics of each neuron, the input they receive and on the connections between neurons. In this article we study the Markovian nature…

应用统计 · 统计学 2012-11-07 Jonathan Touboul , Olivier Faugeras

Markov decision processes (MDPs) are a standard model for sequential decision-making problems and are widely used across many scientific areas, including formal methods and artificial intelligence (AI). MDPs do, however, come with the…

人工智能 · 计算机科学 2024-12-11 Marnix Suilen , Thom Badings , Eline M. Bovy , David Parker , Nils Jansen

We develop an exhaustive study of Markov decision process (MDP) under mean field interaction both on states and actions in the presence of common noise, and when optimization is performed over open-loop controls on infinite horizon. Such…

最优化与控制 · 数学 2021-09-10 Médéric Motte , Huyên Pham

We consider the problem of approximating the reachability probabilities in Markov decision processes (MDP) with uncountable (continuous) state and action spaces. While there are algorithms that, for special classes of such MDP, provide a…

系统与控制 · 电气工程与系统科学 2022-07-13 Kush Grover , Jan Křetínský , Tobias Meggendorfer , Maximilian Weininger

We present a numerical method to compute the survival function and the moments of the exit time for a piecewise-deterministic Markov process (PDMP). Our approach is based on the quantization of an underlying discrete-time Markov chain…

概率论 · 数学 2011-08-31 Adrien Brandejsky , Benoîte de Saporta , François Dufour

Dealing with unichain MDPs, we consider stationary distributions of policies that coincide in all but $n$ states. In these states each policy chooses one of two possible actions. We show that the stationary distributions of n+1 such…

概率论 · 数学 2007-05-23 Ronald Ortner

We study the offline data-driven sequential decision making problem in the framework of Markov decision process (MDP). In order to enhance the generalizability and adaptivity of the learned policy, we propose to evaluate each policy by a…

统计理论 · 数学 2021-11-11 Zhengling Qi , Peng Liao

This paper studies the computation of robust deterministic policies for Markov Decision Processes (MDPs) in the Lightning Does Not Strike Twice (LDST) model of Mannor, Mebel and Xu (ICML '12). In this model, designed to provide robustness…

最优化与控制 · 数学 2024-12-18 Fei Wu , Erik Demeulemeester , Jannik Matuschke

We cast episodic Markov decision process (MDP) planning as Bayesian inference over policies. A policy is treated as the latent variable and is assigned an unnormalized probability of optimality that is monotone in its expected return,…

机器学习 · 计算机科学 2026-04-14 David Tolpin

In this work we investigate an importance sampling approach for evaluating policies for a structurally time-varying factored Markov decision process (MDP), i.e. the policy's value is estimated with a high-probability confidence interval. In…

系统与控制 · 电气工程与系统科学 2023-02-07 Carmel Fiscko , Soummya Kar , Bruno Sinopoli

The mean time taken by an irreducible Markov chain on a finite state space to hit a target chosen at random according to the stationary distribution does not depend on the initial state of the chain. This mean time is known as Kemeny's…

概率论 · 数学 2026-02-13 P. J. Fitzsimmons

We consider non-standard Markov Decision Processes (MDPs) where the target function is not only a simple expectation of the accumulated reward. Instead, we consider rather general functionals of the joint distribution of terminal state and…

最优化与控制 · 数学 2025-10-16 Nicole Bäuerle , Tamara Göll , Anna Jaśkiewicz

Long-run average optimization problems for Markov decision processes (MDPs) require constructing policies with optimal steady-state behavior, i.e., optimal limit frequency of visits to the states. However, such policies may suffer from…

多智能体系统 · 计算机科学 2023-12-20 David Klaška , Antonín Kučera , Vojtěch Kůr , Vít Musil , Vojtěch Řehák

In this paper, we consider a Markov decision process (MDP) with a Borel state space $\textbf{X}\cup\{\Delta\}$, where $\Delta$ is an absorbing state (cemetery), and a Borel action space $\textbf{A}$. We consider the space of finite…

最优化与控制 · 数学 2023-07-07 Alexey Piunovskiy , Yi Zhang

Probabilistic model checking can provide formal guarantees on the behavior of stochastic models relating to a wide range of quantitative properties, such as runtime, energy consumption or cost. But decision making is typically with respect…

计算机科学中的逻辑 · 计算机科学 2024-03-19 Ingy Elsayed-Aly , David Parker , Lu Feng

Ranking the spreading influence of nodes is of great importance in practice and research. The key to ranking a node's spreading ability is to evaluate the fraction of susceptible nodes been infected by the target node during the outbreak,…

物理与社会 · 物理学 2023-02-22 Jian-Hong Lin , Zhao Yang , Jian-Guo Liu , Bo-Lun Chen , Claudio J. Tessone

A semi-Markov process is one that changes states in accordance with a Markov chain but takes a random amount of time between changes. We consider the generalisation to semi-Markov processes of the classical Lamperti law for the occupation…

统计力学 · 物理学 2022-07-13 Théo Dessertaine , Claude Godrèche , Jean-Philippe Bouchaud

We study the problem of synthesizing a policy that maximizes the entropy of a Markov decision process (MDP) subject to a temporal logic constraint. Such a policy minimizes the predictability of the paths it generates, or dually, maximizes…

最优化与控制 · 数学 2019-06-17 Yagiz Savas , Melkior Ornik , Murat Cubuktepe , Mustafa O. Karabag , Ufuk Topcu

For a Markov decision process with countably infinite states, the optimal value may not be achievable in the set of stationary policies. In this paper, we study the existence conditions of an optimal stationary policy in a countable-state…

最优化与控制 · 数学 2020-07-06 Li Xia , Xianping Guo , Xi-Ren Cao