Related papers: Optimal strategies for controlled growth in metast…
We introduce the spatiotemporal Markov decision process (STMDP), a special type of Markov decision process that models sequential decision-making problems which are not only characterized by temporal, but also by spatial interaction…
We investigate the large-time scaling regimes arising from a variety of metastable structures in a chain of Ising spins with both first- and second-neighbor couplings while subject to a Kawasaki dynamics. Depending on the ratio and sign of…
Optimal growth of structures governed by spatially stochastic dynamics arises in many scientific settings, for example in processes such as solution-based crystallization and the formation of microbial biofilms on patterned substrates or…
Interval Markov decision processes are a class of Markov models where the transition probabilities between the states belong to intervals. In this paper, we study the problem of efficient estimation of the optimal policies in Interval…
In this paper, we consider reinforcement learning of Markov Decision Processes (MDP) with peak constraints, where an agent chooses a policy to optimize an objective and at the same time satisfy additional constraints. The agent has to take…
A Markov decision process (MDP) framework is adopted to represent ensemble control of devices with cyclic energy consumption patterns, e.g., thermostatically controlled loads. Specifically we utilize and develop the class of MDP models…
In this paper we analyze the metastable behavior for the Ising model that evolves under Kawasaki dynamics on the hexagonal lattice $\mathbb{H}^2$ in the limit of vanishing temperature. Let $\Lambda\subset\mathbb{H}^2$ a finite set which we…
The standard Markov Decision Process (MDP) formulation hinges on the assumption that an action is executed immediately after it was chosen. However, assuming it is often unrealistic and can lead to catastrophic failures in applications such…
We are interested in risk constraints for infinite horizon discrete time Markov decision processes (MDPs). Starting with average reward MDPs, we show that increasing concave stochastic dominance constraints on the empirical distribution of…
Markov Decision Processes (MDPs) have been used to formulate many decision-making problems in science and engineering. The objective is to synthesize the best decision (action selection) policies to maximize expected rewards (or minimize…
This paper discusses the functional stability of closed-loop Markov Chains under optimal policies resulting from a discounted optimality criterion, forming Markov Decision Processes (MDPs). We investigate the stability of MDPs in the sense…
Optimal control in non-stationary Markov decision processes (MDP) is a challenging problem. The aim in such a control problem is to maximize the long-term discounted reward when the transition dynamics or the reward function can change over…
In this paper we analyze metastability and nucleation in the context of the Kawasaki dynamics for the two-dimensional Ising lattice gas at very low temperature with periodic boundary conditions. Let $\beta>0$ be the inverse temperature and…
It is well known that for any finite state Markov decision process (MDP) there is a memoryless deterministic policy that maximizes the expected reward. For partially observable Markov decision processes (POMDPs), optimal memoryless policies…
This is the second in a series of three papers in which we study a two-dimensional lattice gas consisting of two types of particles subject to Kawasaki dynamics at low temperature in a large finite box with an open boundary. Each pair of…
This paper studies the computation of robust deterministic policies for Markov Decision Processes (MDPs) in the Lightning Does Not Strike Twice (LDST) model of Mannor, Mebel and Xu (ICML '12). In this model, designed to provide robustness…
We consider a ferromagnetic Ising chain evolving under Kawasaki dynamics at zero temperature. We investigate the statistics of the metastable configurations in which the system gets blocked (statistics of energy, spin correlations,…
We present a method for a certain class of Markov Decision Processes (MDPs) that can relate the optimal policy back to one or more reward sources in the environment. For a given initial state, without fully computing the value function,…
Modern large-scale computing deployments consist of complex applications running over machine clusters. An important issue in these is the offering of elasticity, i.e., the dynamic allocation of resources to applications to meet fluctuating…
Markov decision processes (MDPs) are a popular model for performance analysis and optimization of stochastic systems. The parameters of stochastic behavior of MDPs are estimates from empirical observations of a system; their values are not…