English
Related papers

Related papers: On Markov Decision Processes with Borel Spaces and…

200 papers

In this paper we consider the problem of computing an $\epsilon$-optimal policy of a discounted Markov Decision Process (DMDP) provided we can only access its transition function through a generative sampling model that given any…

Optimization and Control · Mathematics 2019-06-07 Aaron Sidford , Mengdi Wang , Xian Wu , Lin F. Yang , Yinyu Ye

We present a general framework for applying machine-learning algorithms to the verification of Markov decision processes (MDPs). The primary goal of these techniques is to improve performance by avoiding an exhaustive exploration of the…

Markov decision processes (MDPs) with multi-dimensional weights are useful to analyze systems with multiple objectives that may be conflicting and require the analysis of trade-offs. We study the complexity of percentile queries in such…

Logic in Computer Science · Computer Science 2016-12-08 Mickael Randour , Jean-François Raskin , Ocan Sankur

We prove new upper and lower bounds for sample complexity of finding an $\epsilon$-optimal policy of an infinite-horizon average-reward Markov decision process (MDP) given access to a generative model. When the mixing time of the…

Machine Learning · Computer Science 2021-06-15 Yujia Jin , Aaron Sidford

We present metrics for measuring state similarity in Markov decision processes (MDPs) with infinitely many states, including MDPs with continuous state spaces. Such metrics provide a stable quantitative analogue of the notion of…

Artificial Intelligence · Computer Science 2012-07-09 Norman Ferns , Prakash Panangaden , Doina Precup

We consider reinforcement learning for continuous-time Markov decision processes (MDPs) in the infinite-horizon, average-reward setting. In contrast to discrete-time MDPs, a continuous-time process moves to a state and stays there for a…

Machine Learning · Computer Science 2024-07-03 Xuefeng Gao , Xun Yu Zhou

We tackle the problem of acting in an unknown finite and discrete Markov Decision Process (MDP) for which the expected shortest path from any state to any other state is bounded by a finite number $D$. An MDP consists of $S$ states and $A$…

Machine Learning · Computer Science 2019-07-11 Aristide Tossou , Christos Dimitrakakis , Debabrota Basu

We consider a novel queuing problem where the decision-maker must choose to accept or reject randomly arriving tasks into a no buffer queue which are processed by $N$ identical servers. Each task has a price, which is a positive real…

Machine Learning · Computer Science 2022-03-16 Marc Rigter , Danial Dervovic , Parisa Hassanzadeh , Jason Long , Parisa Zehtabi , Daniele Magazzeni

The convex analytic method has proved to be a very versatile method for the study of infinite horizon average cost optimal stochastic control problems. In this paper, we revisit the convex analytic method and make three primary…

Optimization and Control · Mathematics 2022-08-04 Ari Arapostathis , Serdar Yüksel

We consider the problem of approximating the reachability probabilities in Markov decision processes (MDP) with uncountable (continuous) state and action spaces. While there are algorithms that, for special classes of such MDP, provide a…

Systems and Control · Electrical Eng. & Systems 2022-07-13 Kush Grover , Jan Křetínský , Tobias Meggendorfer , Maximilian Weininger

We resolve the open question regarding the sample complexity of policy learning for maximizing the long-run average reward associated with a uniformly ergodic Markov decision process (MDP), assuming a generative model. In this context, the…

Machine Learning · Computer Science 2024-02-14 Shengbo Wang , Jose Blanchet , Peter Glynn

Value-at-risk (VaR), also known as quantile, is a crucial risk measure in finance and other fields. However, optimizing VaR metrics in Markov decision processes (MDPs) is challenging because VaR is non-additive and the traditional dynamic…

Optimization and Control · Mathematics 2025-07-31 Li Xia , Jinyan Pan

We prove that the simplex method with the highest gain/most-negative-reduced cost pivoting rule converges in strongly polynomial time for deterministic Markov decision processes (MDPs) regardless of the discount factor. For a deterministic…

Data Structures and Algorithms · Computer Science 2013-02-01 Ian Post , Yinyu Ye

We consider Markov Decision Processes (MDPs) in which every stationary policy induces the same graph structure for the underlying Markov chain and further, the graph has the following property: if we replace each recurrent class by a node,…

Machine Learning · Computer Science 2021-03-10 Joseph Lubars , Anna Winnicki , Michael Livesay , R. Srikant

The standard version of the policy iteration (PI) algorithm fails for semicontinuous models, that is, for models with lower semicontinuous one-step costs and weakly continuous transition law. This is due to the lack of continuity properties…

Optimization and Control · Mathematics 2023-07-17 Óscar Vega-Amaya , Fernando Luque-Vásquez

Energy Markov Decision Processes (EMDPs) are finite-state Markov decision processes where each transition is assigned an integer counter update and a rational payoff. An EMDP configuration is a pair s(n), where s is a control state and n is…

Logic in Computer Science · Computer Science 2016-07-05 Tomáš Brázdil , Antonín Kučera , Petr Novotný

A new approach to computation of optimal policies for MDP (Markov decision process) models is introduced. The main idea is to solve not one, but an entire family of MDPs, parameterized by a weighting factor $\zeta$ that appears in the…

Optimization and Control · Mathematics 2018-09-18 Ana Bušić , Sean Meyn

We consider statistical Markov Decision Processes where the decision maker is risk averse against model ambiguity. The latter is given by an unknown parameter which influences the transition law and the cost functions. Risk aversion is…

Optimization and Control · Mathematics 2021-07-21 Nicole Bäuerle , Ulrich Rieder

We study Markov decision processes with Polish state and action spaces. The action space is state dependent and is not necessarily compact. We first establish the existence of an optimal ergodic occupation measure using only a near-monotone…

Optimization and Control · Mathematics 2023-08-15 Ari Arapostathis , Vivek S. Borkar

The average cost optimality is known to be a challenging problem for partially observable stochastic control, with few results available beyond the finite state, action, and measurement setup, for which somewhat restrictive conditions are…

Optimization and Control · Mathematics 2024-08-01 Yunus Emre Demirci , Ali Devran Kara , Serdar Yüksel