中文
相关论文

相关论文: Methods for computing state similarity in Markov D…

200 篇论文

In this paper we consider the problem of computing an $\epsilon$-optimal policy of a discounted Markov Decision Process (DMDP) provided we can only access its transition function through a generative sampling model that given any…

最优化与控制 · 数学 2019-06-07 Aaron Sidford , Mengdi Wang , Xian Wu , Lin F. Yang , Yinyu Ye

This paper considers the problem of remote state estimation for Markov jump linear systems in the presence of uncertainty in the posterior mode probabilities. Such uncertainty may arise when the estimator receives noisy or incomplete…

系统与控制 · 电气工程与系统科学 2025-09-05 Ioannis Tzortzis , Themistoklis Charalambous , Charalambos D. Charalambous

In many operations management problems, we need to make decisions sequentially to minimize the cost while satisfying certain constraints. One modeling approach to study such problems is constrained Markov decision process (CMDP). When…

最优化与控制 · 数学 2021-01-27 Yi Chen , Jing Dong , Zhaoran Wang

In this paper, we show the convergence rates of posterior distributions of the model dynamics in a MDP for both episodic and continuous tasks. The theoretical results hold for general state and action space and the parameter space of the…

统计理论 · 数学 2019-07-23 Zhen Li , Eric Laber

This paper proposes a new probabilistic classification algorithm using a Markov random field approach. The joint distribution of class labels is explicitly modelled using the distances between feature vectors. Intuitively, a class label…

统计计算 · 统计学 2010-06-02 Nial Friel , Anthony N. Pettitt

We study upper and lower bounds on the sample-complexity of learning near-optimal behaviour in finite-state discounted Markov Decision Processes (MDPs). For the upper bound we make the assumption that each action leads to at most two…

机器学习 · 计算机科学 2013-05-17 Tor Lattimore , Marcus Hutter

Markov decision processes (MDP) are useful to model optimisation problems in concurrent systems. To verify MDPs with efficient Monte Carlo techniques requires that their nondeterminism be resolved by a scheduler. Recent work has introduced…

数据结构与算法 · 计算机科学 2016-11-15 Pedro D'Argenio , Axel Legay , Sean Sedwards , Louis-Marie Traonouez

Markov Chain Monte Carlo (MCMC) methods have a drawback when working with a target distribution or likelihood function that is computationally expensive to evaluate, specially when working with big data. This paper focuses on…

机器学习 · 计算机科学 2019-10-22 Asif J. Chowdhury , Gabriel Terejanu

Markov Decision Processes (MDPs) are a popular class of models suitable for solving control decision problems in probabilistic reactive systems. We consider parametric MDPs (pMDPs) that include parameters in some of the transition…

计算机科学中的逻辑 · 计算机科学 2018-06-14 Sebastian Arming , Ezio Bartocci , Krishnendu Chatterjee , Joost-Pieter Katoen , Ana Sokolova

Piecewise-Deterministic Markov Processes (PDMPs) hold significant promise for sampling from complex probability distributions. However, their practical implementation is hindered by the need to compute model-specific bounds. Conversely,…

统计计算 · 统计学 2025-03-17 Augustin Chevallier , Sam Power , Matthew Sutton

Distances between quantum states are reviewed within the framework of the tomographic-probability representation. Tomographic approach is based on observed probabilities and is straightforward for data processing. Different states are…

量子物理 · 物理学 2010-10-12 S. N. Filippov , V. I. Man'ko

Bayesian inference in state-space models is challenging due to high-dimensional state trajectories. A viable approach is particle Markov chain Monte Carlo, combining MCMC and sequential Monte Carlo to form "exact approximations" to…

统计计算 · 统计学 2022-10-27 Anna Wigren , Riccardo Sven Risuleo , Lawrence Murray , Fredrik Lindsten

A quantum ensemble $\{(p_x, \rho_x)\}$ is a set of quantum states each occurring randomly with a given probability. Quantum ensembles are necessary to describe situations with incomplete a priori information, such as the output of a…

量子物理 · 物理学 2009-03-30 Ognyan Oreshkov , John Calsamiglia

We propose a new approach to the problem of searching a space of policies for a Markov decision process (MDP) or a partially observable Markov decision process (POMDP), given a model. Our approach is based on the following observation: Any…

人工智能 · 计算机科学 2013-01-18 Andrew Y. Ng , Michael I. Jordan

We study computational and statistical aspects of learning Latent Markov Decision Processes (LMDPs). In this model, the learner interacts with an MDP drawn at the beginning of each epoch from an unknown mixture of MDPs. To sidestep known…

机器学习 · 计算机科学 2024-06-13 Fan Chen , Constantinos Daskalakis , Noah Golowich , Alexander Rakhlin

Much recent research in decision theoretic planning has adopted Markov decision processes (MDPs) as the model of choice, and has attempted to make their solution more tractable by exploiting problem structure. One particular algorithm,…

人工智能 · 计算机科学 2013-02-08 Craig Boutilier

We consider reinforcement learning in parameterized Markov Decision Processes (MDPs), where the parameterization may induce correlation across transition probabilities or rewards. Consequently, observing a particular state transition might…

机器学习 · 统计学 2015-04-01 Aditya Gopalan , Shie Mannor

Many problems in sequential decision making and stochastic control often have natural multiscale structure: sub-tasks are assembled together to accomplish complex goals. Systematically inferring and leveraging hierarchical structure,…

人工智能 · 计算机科学 2012-12-06 Jake Bouvrie , Mauro Maggioni

Several researchers have proposed minimisation of maximum mean discrepancy (MMD) as a method to quantise probability measures, i.e., to approximate a target distribution by a representative point set. We consider sequential algorithms that…

机器学习 · 统计学 2021-02-15 Onur Teymur , Jackson Gorham , Marina Riabiz , Chris. J. Oates

In this paper, we analyze the convergence behavior of Hermite-type sampling Kantorovich operators in the context of mixed norm spaces. We prove certain direct approximation theorems, including the uniform convergence theorem, the…

泛函分析 · 数学 2025-06-04 Puja Sonawane , A. Sathish Kumar
‹ 上一页 1 8 9 10 下一页 ›