中文
相关论文

相关论文: Methods for computing state similarity in Markov D…

200 篇论文

This paper presents a state representation for reward-free Markov decision processes. The idea is to learn, in a self-supervised manner, an embedding space where distances between pairs of embedded states correspond to the minimum number of…

机器学习 · 计算机科学 2023-12-20 Lorenzo Steccanella , Anders Jonsson

We consider the problem belief-state monitoring for the purposes of implementing a policy for a partially-observable Markov decision process (POMDP), specifically how one might approximate the belief state. Other schemes for belief-state…

人工智能 · 计算机科学 2013-01-18 Pascal Poupart , Craig Boutilier

We investigate the use of temporally abstract actions, or macro-actions, in the solution of Markov decision processes. Unlike current models that combine both primitive actions and macro-actions and leave the state space unchanged, we…

人工智能 · 计算机科学 2013-02-01 Milos Hauskrecht , Nicolas Meuleau , Leslie Pack Kaelbling , Thomas L. Dean , Craig Boutilier

Multistate dynamical processes on networks, where nodes can occupy one of a multitude of discrete states, are gaining widespread use because of their ability to recreate realistic, complex behaviour that cannot be adequately captured by…

物理与社会 · 物理学 2017-09-29 Peter G. Fennell , James P. Gleeson

The task of state estimation in active distribution systems faces a major challenge due to the integration of different measurements with multiple reporting rates. As a result, distribution systems are essentially unobservable in real time,…

最优化与控制 · 数学 2024-05-13 J. G. De la Varga , S. Pineda , J. M. Morales , Á. Porras

To sample from a given target distribution, Markov chain Monte Carlo (MCMC) sampling relies on constructing an ergodic Markov chain with the target distribution as its invariant measure. For any MCMC method, an important question is how to…

概率论 · 数学 2023-08-15 Federica Milinanni , Pierre Nyquist

Reachability analysis of hybrid systems has been used as a safety verification tool to assess offline whether the state of a system is capable of remaining within a designated safe region for a given time horizon. Although it has been…

最优化与控制 · 数学 2014-04-24 Kendra Lesser , Meeko Oishi

The goal of this thesis is to study the use of the Kantorovich-Rubinstein distance as to build a descriptor of sample complexity in classification problems. The idea is to use the fact that the Kantorovich-Rubinstein distance is a metric in…

概率论 · 数学 2023-09-19 Gaël Giordano

The approximation of a discrete probability distribution $\mathbf{t}$ by an $M$-type distribution $\mathbf{p}$ is considered. The approximation error is measured by the informational divergence $\mathbb{D}(\mathbf{t}\Vert\mathbf{p})$, which…

信息论 · 计算机科学 2016-07-28 Bernhard C. Geiger , Georg Böcherer

This paper extends to Continuous-Time Jump Markov Decision Processes (CTJMDP) the classic result for Markov Decision Processes stating that, for a given initial state distribution, for every policy there is a (randomized) Markov policy,…

最优化与控制 · 数学 2020-05-18 Eugene A. Feinberg , Manasa Mandava , Albert N. Shiryaev

Many studies involving large Markov chains require determining a smaller representative (aggregated) chains. Each {\em superstate} in the representative chain represents a {\em group of related} states in the original Markov chain.…

系统与控制 · 电气工程与系统科学 2021-02-19 Amber Srivastava , Raj K. Velicheti , Srinivasa M. Salapaka

This paper studies discrete-time average-cost infinite-horizon Markov decision processes (MDPs) with Borel state and action sets. It introduces new sufficient conditions for { the} validity of optimality inequalities and optimality…

最优化与控制 · 数学 2025-01-28 Eugene A. Feinberg , Pavlo O. Kasyanov , Liliia S. Paliichuk

A Markov decision process (MDP) framework is adopted to represent ensemble control of devices with cyclic energy consumption patterns, e.g., thermostatically controlled loads. Specifically we utilize and develop the class of MDP models…

系统与控制 · 计算机科学 2017-10-24 Michael Chertkov , Vladimir Y. Chernyak , Deepjyoti Deka

The goal of this paper is to analyze distributional Markov Decision Processes as a class of control problems in which the objective is to learn policies that steer the distribution of a cumulative reward toward a prescribed target law,…

最优化与控制 · 数学 2026-02-09 Nicole Bäuerle , Athanasios Vasileiadis

Markov state models (MSMs) have become a popular approach for investigating the conformational dynamics of proteins and other biomolecules. MSMs are typically built from numerous molecular dynamics simulations by dividing the sampled…

生物大分子 · 定量生物学 2015-06-12 Yuan Yao , Raymond Z. Cui , Gregory R. Bowman , Daniel Silva , Jian Sun , Xuhui Huang

This paper addresses the problem of planning under uncertainty in large Markov Decision Processes (MDPs). Factored MDPs represent a complex state space using state variables and the transition model using a dynamic Bayesian network. This…

人工智能 · 计算机科学 2011-06-10 C. Guestrin , D. Koller , R. Parr , S. Venkataraman

Abstraction of Markov Decision Processes is a useful tool for solving complex problems, as it can ignore unimportant aspects of an environment, simplifying the process of learning an optimal policy. In this paper, we propose a new algorithm…

机器学习 · 计算机科学 2021-04-20 Ondrej Biza , Robert Platt

We introduce $(\varepsilon, \delta)$-bisimulation, a novel type of approximate probabilistic bisimulation for continuous-time Markov chains. In contrast to related notions, $(\varepsilon, \delta)$-bisimulation allows the use of different…

计算机科学中的逻辑 · 计算机科学 2025-05-23 Timm Spork , Christel Baier , Joost-Pieter Katoen , Sascha Klüppelholz , Jakob Piribauer

We consider numerical schemes for computing the linear response of steady-state averages of stochastic dynamics with respect to a perturbation of the drift part of the stochastic differential equation. The schemes are based on Girsanov's…

数值分析 · 数学 2019-12-18 Petr Plechac , Gabriel Stoltz , Ting Wang

We study minority games in efficient regime. By incorporating the utility function and aggregating agents with similar strategies we develop an effective mesoscale notion of state of the game. Using this approach, the game can be…

适应与自组织系统 · 物理学 2011-12-06 Karol Wawrzyniak , Wojciech Wislicki