Related papers: Asymptotics of impulse control problem with multip…
This article proposes an approach to construct a Lyapunov function for a linear coupled impulsive system consisting of two time-invariant subsystems. In contrast to various variants of small-gain stability conditions for coupled systems,…
Navigating a collision-free and optimal trajectory for a robot is a challenging task, particularly in environments with moving obstacles such as humans. We formulate this problem as a stochastic optimal control problem. Since solving the…
The solution to the infinite horizon optimal control problem for linear distributed time-delay systems is presented. The proposal is based on the use of the Cauchy solution for distributed time-delay systems. In contrast with previous…
We study the large time behavior of solutions to fully nonlinear parabolic equations of Hamilton-Jacobi-Bellman type arising typically in stochastic control theory with control both on drift and diffusion coefficients. We prove that, as…
We consider partially observable Markov decision processes (POMDPs) with limit-average payoff, where a reward value in the interval [0,1] is associated to every transition, and the payoff of an infinite path is the long-run average of the…
We consider a control system with dynamics which are affine in the (unbounded) derivative of the control $u$. We introduce a notion of generalized solution $x$ on $[0,T]$ for controls $u$ of bounded total variation on $[0,t]$ for every…
In this paper, motivated by a problem in stochastic impulse control theory, we aim to study solutions to a free boundary problem of obstacle-type. We obtain sharp estimates for the solution using nonlinear tools which are independent of the…
In this paper, we consider a class of stochastic impulse control problem when there is a fixed delay $\Delta$ between the decision and execution times. The dynamics of the controlled system between two impulses is an arbitrary adapted…
We consider a discrete-time Markov decision process with Borel state and action spaces. The performance criterion is to maximize a total expected {utility determined by unbounded return function. It is shown the existence of optimal…
We present for the first time an asymptotic convergence analysis of two time-scale stochastic approximation driven by "controlled" Markov noise. In particular, the faster and slower recursions have non-additive controlled Markov noise…
We study Markov decision processes with Polish state and action spaces. The action space is state dependent and is not necessarily compact. We first establish the existence of an optimal ergodic occupation measure using only a near-monotone…
This paper develops an inverse reinforcement learning algorithm aimed at recovering a reward function from the observed actions of an agent. We introduce a strategy to flexibly handle different types of actions with two approximations of…
The aim of this study is to devise numerical methods for dealing with very high-dimensional Bermudan-style derivatives. For such problems, we quickly see that we can at best hope for price bounds, and we can only use a simulation approach.…
We consider partially observable Markov decision processes (POMDPs) with limit-average payoff, where a reward value in the interval [0,1] is associated to every transition, and the payoff of an infinite path is the long-run average of the…
Under the expected total reward criterion, the optimal value of a finite-horizon Markov decision process can be determined by solving the Bellman equations. The equations were extended by D. J. White to processes with vector rewards in…
An important step in the Markov reward approach to error bounds on stationary performance measures of Markov chains is to bound the bias terms. Affine functions have been successfully used for these bounds for various models, but there are…
We consider the model selection problem for a large class of time series models, including, multivariate count processes, causal processes with exogenous covariates. A procedure based on a general penalized contrast is proposed. Some…
We consider a general type of non-Markovian impulse control problems under adverse non-linear expectation or, more specifically, the zero-sum game problem where the adversary player decides the probability measure. We show that the upper…
Active inference is a probabilistic framework for modelling the behaviour of biological and artificial agents, which derives from the principle of minimising free energy. In recent years, this framework has successfully been applied to a…
We consider a nonlinear control system depending on two controls u and v, with dynamics affine in the (unbounded) derivative of u, and v appearing initially only in the drift term. Recently, motivated by applications to optimization…