Related papers: Asymptotics of impulse control problem with multip…
A Markov decision process can be parameterized by a transition kernel and a reward function. Both play essential roles in the study of reinforcement learning as evidenced by their presence in the Bellman equations. In our inquiry of various…
We consider a simple control problem in which the underlying dynamics depend on a parameter that is unknown and must be learned. We exhibit a control strategy which is optimal to within a multiplicative constant. While most authors find…
This paper presents a class of Dynamic Multi-Armed Bandit problems where the reward can be modeled as the noisy output of a time varying linear stochastic dynamic system that satisfies some boundedness constraints. The class allows many…
We study infinite horizon control of continuous-time non-linear branching processes with almost sure extinction for general (positive or negative) discount. Our main goal is to study the link between infinite horizon control of these…
We study infinite-horizon stochastic optimal control problems with observable side information: a Markov chain that modulates an unknown context-conditional randomness distribution. Since this distribution is unknown, we propose a Bayesian…
Markov decision processes (MDPs) are standard models for probabilistic systems with non-deterministic behaviours. Long-run average rewards provide a mathematically elegant formalism for expressing long term performance. Value iteration (VI)…
We consider stochastic impulse control problems where the process is driven by a general one-dimensional diffusion. We shall show a new mathematical characterization of the value function as a linear function in a certain transformed space.…
For a multivariate random walk with i.i.d. jumps satisfying the Cramer moment condition and having a mean vector with at least one negative component, we derive the exact asymptotics of the probability of ever hitting the positive orthant…
The interaction between an artificial agent and its environment is bi-directional. The agent extracts relevant information from the environment, and affects the environment by its actions in return to accumulate high expected reward.…
In this paper we investigate the long time behavior of solutions to fractional in time evolution equations which appear as results of random time changes in Markov processes. We consider inverse subordinators as random times and use the…
In this paper, we address a social planner's optimal control problem for a partially observable stochastic epidemic model. The control measures include social distancing, testing, and vaccination. Using a diffusion approximation for the…
In this paper, we present a numerical framework for constructing bounds on stationary performance measures of random walks in the positive orthant using the Markov reward approach. These bounds are established in terms of stationary…
This paper, the second of a two-part series, presents a method for mean-field feedback stabilization of a swarm of agents on a finite state space whose time evolution is modeled as a continuous time Markov chain (CTMC). The resulting…
In this paper we consider the optimal control of Hilbert space-valued infinite-dimensional Piecewise Deterministic Markov Processes (PDMP) and we prove that the corresponding value function can be represented via a Feynman-Kac type formula…
In this paper, we propose a new policy iteration algorithm to compute the value function and the optimal controls of continuous time stochastic control problems. The algorithm relies on successive approximations using linear-quadratic…
Deterministic optimal impulse control problem with terminal state constraint is considered. Due to the appearance of the terminal state constraint, the value function might be discontinuous in general. The main contribution of this paper is…
We prove a Carleman estimate for a one-dimensional parabolic equation which degenerates at one extremity of the domain and has a bounded, time dependent coefficient multiplying the diffusion term. Then we use the estimate to show the null…
We consider a system of interacting particles governed by the generalized Langevin equation (GLE) in the presence of external confining potentials, singular repulsive forces, as well as memory kernels. Using a Mori-Zwanzig approach, we…
Planning problems where effects of actions are non-deterministic can be modeled as Markov decision processes. Planning problems are usually goal-directed. This paper proposes several techniques for exploiting the goal-directedness to…
We consider a class of exit--time control problems for nonlinear systems with a nonnegative vanishing Lagrangian. In general, the associated PDE may have multiple solutions, and known regularity and stability properties do not hold. In this…