English
Related papers

Related papers: Asymptotics of impulse control problem with multip…

200 papers

A Markov decision process can be parameterized by a transition kernel and a reward function. Both play essential roles in the study of reinforcement learning as evidenced by their presence in the Bellman equations. In our inquiry of various…

Machine Learning · Computer Science 2023-09-04 Falcon Z. Dai

We consider a simple control problem in which the underlying dynamics depend on a parameter that is unknown and must be learned. We exhibit a control strategy which is optimal to within a multiplicative constant. While most authors find…

Optimization and Control · Mathematics 2021-09-15 Jacob Carruth , Maximilian F. Eggl , Charles Fefferman , Clarence W. Rowley , Melanie Weber

This paper presents a class of Dynamic Multi-Armed Bandit problems where the reward can be modeled as the noisy output of a time varying linear stochastic dynamic system that satisfies some boundedness constraints. The class allows many…

Machine Learning · Computer Science 2017-10-10 T. W. U. Madhushani , D. H. S. Maithripala , N. E. Leonard

We study infinite horizon control of continuous-time non-linear branching processes with almost sure extinction for general (positive or negative) discount. Our main goal is to study the link between infinite horizon control of these…

Probability · Mathematics 2016-07-28 Julien Claisse , Nicolas Champagnat

We study infinite-horizon stochastic optimal control problems with observable side information: a Markov chain that modulates an unknown context-conditional randomness distribution. Since this distribution is unknown, we propose a Bayesian…

Optimization and Control · Mathematics 2026-02-26 Johannes Milz , Alexander Shapiro , Enlu Zhou

Markov decision processes (MDPs) are standard models for probabilistic systems with non-deterministic behaviours. Long-run average rewards provide a mathematically elegant formalism for expressing long term performance. Value iteration (VI)…

Systems and Control · Computer Science 2017-09-01 Pranav Ashok , Krishnendu Chatterjee , Przemyslaw Daca , Jan Křetínský , Tobias Meggendorfer

We consider stochastic impulse control problems where the process is driven by a general one-dimensional diffusion. We shall show a new mathematical characterization of the value function as a linear function in a certain transformed space.…

Optimization and Control · Mathematics 2007-05-23 Masahiko Egami

For a multivariate random walk with i.i.d. jumps satisfying the Cramer moment condition and having a mean vector with at least one negative component, we derive the exact asymptotics of the probability of ever hitting the positive orthant…

Probability · Mathematics 2019-05-09 Yuqing Pan , Konstantin Borovkov

The interaction between an artificial agent and its environment is bi-directional. The agent extracts relevant information from the environment, and affects the environment by its actions in return to accumulate high expected reward.…

Systems and Control · Computer Science 2018-06-06 Stas Tiomkin , Naftali Tishby

In this paper we investigate the long time behavior of solutions to fractional in time evolution equations which appear as results of random time changes in Markov processes. We consider inverse subordinators as random times and use the…

Probability · Mathematics 2020-06-25 Anatoly N. Kochubei , Yuri Kondratiev , José L. da Silva

In this paper, we address a social planner's optimal control problem for a partially observable stochastic epidemic model. The control measures include social distancing, testing, and vaccination. Using a diffusion approximation for the…

Optimization and Control · Mathematics 2025-03-11 Ibrahim Mbouandi Njiasse , Florent Ouabo Kamkumo , Ralf Wunderlich

In this paper, we present a numerical framework for constructing bounds on stationary performance measures of random walks in the positive orthant using the Markov reward approach. These bounds are established in terms of stationary…

Probability · Mathematics 2018-11-22 Xinwei Bai , Jasper Goseling

This paper, the second of a two-part series, presents a method for mean-field feedback stabilization of a swarm of agents on a finite state space whose time evolution is modeled as a continuous time Markov chain (CTMC). The resulting…

Systems and Control · Computer Science 2017-03-29 Shiba Biswal , Karthik Elamvazhuthi , Spring Berman

In this paper we consider the optimal control of Hilbert space-valued infinite-dimensional Piecewise Deterministic Markov Processes (PDMP) and we prove that the corresponding value function can be represented via a Feynman-Kac type formula…

Optimization and Control · Mathematics 2019-06-07 Elena Bandini , Michele Thieullen

In this paper, we propose a new policy iteration algorithm to compute the value function and the optimal controls of continuous time stochastic control problems. The algorithm relies on successive approximations using linear-quadratic…

Optimization and Control · Mathematics 2024-09-09 Dylan Possamaï , Ludovic Tangpi

Deterministic optimal impulse control problem with terminal state constraint is considered. Due to the appearance of the terminal state constraint, the value function might be discontinuous in general. The main contribution of this paper is…

Optimization and Control · Mathematics 2020-11-10 Yue Zhou , Xinwei Feng , Jiongmin Yong

We prove a Carleman estimate for a one-dimensional parabolic equation which degenerates at one extremity of the domain and has a bounded, time dependent coefficient multiplying the diffusion term. Then we use the estimate to show the null…

Analysis of PDEs · Mathematics 2025-08-26 Alfredo S. Gamboa , Juan Limaco , Luis P. Yapu

We consider a system of interacting particles governed by the generalized Langevin equation (GLE) in the presence of external confining potentials, singular repulsive forces, as well as memory kernels. Using a Mori-Zwanzig approach, we…

Probability · Mathematics 2024-03-15 Manh Hong Duong , Hung D. Nguyen

Planning problems where effects of actions are non-deterministic can be modeled as Markov decision processes. Planning problems are usually goal-directed. This paper proposes several techniques for exploiting the goal-directedness to…

Artificial Intelligence · Computer Science 2013-02-08 Nevin Lianwen Zhang , Weihong Zhang

We consider a class of exit--time control problems for nonlinear systems with a nonnegative vanishing Lagrangian. In general, the associated PDE may have multiple solutions, and known regularity and stability properties do not hold. In this…

Optimization and Control · Mathematics 2018-05-10 Monica Motta , Caterina Sartori