中文
相关论文

相关论文: Transient Reward Approximation for Continuous-Time…

200 篇论文

Reinforcement Learning (RL) has gained substantial attention across diverse application domains and theoretical investigations. Existing literature on RL theory largely focuses on risk-neutral settings where the decision-maker learns to…

机器学习 · 计算机科学 2024-12-24 Zhengqi Wu , Renyuan Xu

Pairwise Choice Markov Chains (PCMC) have been recently introduced to overcome limitations of choice models based on traditional axioms unable to express empirical observations from modern behavior economics like context effects occurring…

机器学习 · 计算机科学 2020-02-03 Alix Lhéritier

This paper presents an efficient procedure for multi-objective model checking of long-run average reward (aka: mean pay-off) and total reward objectives as well as their combination. We consider this for Markov automata, a compositional…

计算机科学中的逻辑 · 计算机科学 2021-01-08 Tim Quatmann , Joost-Pieter Katoen

Constrained partially observable Markov decision processes (CPOMDPs) have been used to model various real-world phenomena. However, they are notoriously difficult to solve to optimality, and there exist only a few approximation methods for…

人工智能 · 计算机科学 2023-06-27 Robert K. Helmeczi , Can Kavaklioglu , Mucahit Cevik

The general sequential decision-making problem, which includes Markov decision processes (MDPs) and partially observable MDPs (POMDPs) as special cases, aims at maximizing a cumulative reward by making a sequence of decisions based on a…

机器学习 · 计算机科学 2024-02-07 Ruiquan Huang , Yingbin Liang , Jing Yang

We are interested in risk constraints for infinite horizon discrete time Markov decision processes (MDPs). Starting with average reward MDPs, we show that increasing concave stochastic dominance constraints on the empirical distribution of…

最优化与控制 · 数学 2012-06-21 William B. Haskell , Rahul Jain

While there is an extensive body of research on the analysis of Value Iteration (VI) for discounted cumulative-reward MDPs, prior work on analyzing VI for (undiscounted) average-reward MDPs has been limited, and most prior results focus on…

最优化与控制 · 数学 2026-02-10 Jongmin Lee , Ernest K. Ryu

Although the Bayesian paradigm offers a formal framework for estimating the entire probability distribution over uncertain parameters, its online implementation can be challenging due to high computational costs. We suggest the Adaptive…

机器学习 · 计算机科学 2023-10-23 Pedram Agand , Mo Chen , Hamid D. Taghirad

Reversible jump Markov chain Monte Carlo (RJMCMC) proposals that achieve reasonable acceptance rates and mixing are notoriously difficult to design in most applications. Inspired by recent advances in deep neural network-based normalizing…

统计计算 · 统计学 2023-02-28 Laurence Davies , Robert Salomone , Matthew Sutton , Christopher Drovandi

In this paper, we develop a general law of large numbers and central limit theorem for cumulative reward processes associated with finite state Markov jump processes with non-stationary transition rates. Such models commonly arise in…

概率论 · 数学 2025-10-15 Monte Fischer , Peter W. Glynn

In this paper, we study a mean-variance optimization problem in an infinite horizon discrete time discounted Markov decision process (MDP). The objective is to minimize the variance of system rewards with the constraint of mean performance.…

最优化与控制 · 数学 2017-08-24 Li Xia

In this research paper, weighted / unweighted, directed / undirected graphs are associated with interesting Discrete Time Markov Chains (DTMCs) as well as Continuous Time Markov Chains (CTMCs). The equilibrium / transient behaviour of such…

数据结构与算法 · 计算机科学 2012-09-18 Garimella Rama Murthy

We study a quantum entanglement switch that serves $k$ users in a star topology. We model variants of the system using Markov chains and standard queueing theory and obtain expressions for switch capacity and the expected number of qubits…

量子物理 · 物理学 2022-04-27 Gayane Vardoyan , Saikat Guha , Philippe Nain , Don Towsley

Motivated by robotic surveillance applications, this paper studies the novel problem of maximizing the return time entropy of a Markov chain, subject to a graph topology with travel times and stationary distribution. The return time entropy…

最优化与控制 · 数学 2018-05-29 Xiaoming Duan , Mishel George , Francesco Bullo

In robust Markov decision processes (RMDPs), it is assumed that the reward and the transition dynamics lie in a given uncertainty set. By targeting maximal return under the most adversarial model from that set, RMDPs address performance…

机器学习 · 计算机科学 2024-02-13 Uri Gadot , Esther Derman , Navdeep Kumar , Maxence Mohamed Elfatihi , Kfir Levy , Shie Mannor

Performing numerical integration when the integrand itself cannot be evaluated point-wise is a challenging task that arises in statistical analysis, notably in Bayesian inference for models with intractable likelihood functions. Markov…

统计计算 · 统计学 2020-06-17 Lawrence Middleton , George Deligiannidis , Arnaud Doucet , Pierre E. Jacob

Recent advances have significantly improved our understanding of the sample complexity of learning in average-reward Markov decision processes (AMDPs) under the generative model. However, much less is known about the constrained…

机器学习 · 计算机科学 2025-09-23 Yukuan Wei , Xudong Li , Lin F. Yang

We introduce causal Markov Decision Processes (C-MDPs), a new formalism for sequential decision making which combines the standard MDP formulation with causal structures over state transition and reward functions. Many contemporary and…

机器学习 · 统计学 2021-02-16 Yangyi Lu , Amirhossein Meisami , Ambuj Tewari

We study the verification of a finite continuous-time Markov chain (CTMC) C against a linear real-time specification given as a deterministic timed automaton (DTA) A with finite or Muller acceptance conditions. The central question that we…

计算机科学中的逻辑 · 计算机科学 2015-07-01 Taolue Chen , Tingting Han , Joost-Pieter Katoen , Alexandru Mereacre

There are no computationally feasible algorithms that provide solutions to the finite horizon Risk-sensitive Constrained Markov Decision Process (Risk-CMDP) problem, even for problems with moderate horizon. With an aim to design the same,…

最优化与控制 · 数学 2023-03-27 Vartika Singh , Veeraruna Kavitha