Related papers: Resolvent-Techniques For Multiple Exercise Problem…
In many engineered systems, optimization is used for decision making at time-scales ranging from real-time operation to long-term planning. This process often involves solving similar optimization problems over and over again with slightly…
In the literature on optimal stopping, the problem of maximizing the expected discounted reward over all stopping times has been explicitly solved for some special reward functions (including $(x^+)^{\nu}$, $(e^x-K)^+$, $(K-e^{-x})^+$,…
Many potential applications of reinforcement learning (RL) require guarantees that the agent will perform well in the face of disturbances to the dynamics or reward function. In this paper, we prove theoretically that maximum entropy…
In this paper, the reinforcement learning (RL)-based optimal control problem is studied for multiplicative-noise systems, where input delay is involved and partial system dynamics is unknown. To solve a variant of Riccati-ZXL equations,…
The purpose of this paper is to consider the exit-time problem for a finite-range Markov jump process, i.e, the distance the particle can jump is bounded independent of its location. Such jump diffusions are expedient models for anomalous…
In this paper, we study the asymptotic of exit problem for controlled Markov diffusion processes with random jumps and vanishing diffusion terms, where the random jumps are introduced in order to modify the evolution of the controlled…
We consider the problem of computing optimal policies in average-reward Markov decision processes. This classical problem can be formulated as a linear program directly amenable to saddle-point optimization methods, albeit with a number of…
We study the optimal stopping problem of McKean-Vlasov diffusions when the criterion is a function of the law of the stopped process. A remarkable new feature in this setting is that the stopping time also impacts the dynamics of the…
This paper is devoted to studying the average optimality in continuous-time Markov decision processes with fairly general state and action spaces. The criterion to be maximized is expected average rewards. The transition rates of underlying…
Via operator theoretic methods, we formalize the concentration phenomenon for a given observable `$r$' of a discrete time Markov chain with `$\mu_{\pi}$' as invariant ergodic measure, possibly having support on an unbounded state space. The…
We investigate the optimal stopping problems involving the supremum of a diffusion. The starting point is the link between works of Peskir and Meilijson, which we describe in a unified manner. The description developped follows mainly the…
We study the local regularity and multifractal nature of the sample paths of jump diffusion processes, which are solutions to a class of stochastic differential equations with jumps. This article extends the recent work of Barral {\it et…
Random tensors can be used to produce random matrices. This idea is, for instance, very natural when one studies random quantum states with the aim of exploring properties that are generically true, or true with some probability. We hereby…
Let $D\subset R^d$ be a bounded domain and denote by $\mathcal P(D)$ the space of probability measures on $D$. Let \begin{equation*} L=\frac12\nabla\cdot a\nabla +b\nabla \end{equation*} be a second order elliptic operator. Let…
This paper studies the continuous-time reinforcement learning (RL) for optimal switching problems across multiple regimes. We consider a type of exploratory formulation under entropy regularization where the agent randomizes both the timing…
We extend to multi-dimensions the work of [1], where new fully explicit kinetic methods were built for the approximation of linear and non-linear convection-diffusion problems. The fundamental principles from the earlier work are retained:…
We use the geometry of suitably generalised potentials to solve risk-sensitive Markovian optimal stopping problems. As in the linear case due to Dynkin and Yushkievich (1967), the value function is the pointwise infimum of those functions…
We provide resolvent asymptotics as well as various operator-norm estimates for the system of linear partial differential equations describing the thin infinite elastic rod with material coefficients which periodically highly oscillate…
By the recent advances in computer technology leading to the invention of more powerful processors, the importance of creating models using data training is even greater than ever. Given the significance of this issue, this work tries to…
This paper considers optimization over multiple renewal systems coupled by time average constraints. These systems act asynchronously over variable length frames. For each system, at the beginning of each renewal frame, it chooses an action…