Related papers: An Entropy Regularized BSDE Approach to Bermudan O…
We study automated intrusion prevention using reinforcement learning. In a novel approach, we formulate the problem of intrusion prevention as an optimal stopping problem. This formulation allows us insight into the structure of the optimal…
In this paper we introduce and study the concept of optimal and surely optimal dual martingales in the context of dual valuation of Bermudan options, and outline the development of new algorithms in this context. We provide a…
The aim of this study is to devise numerical methods for dealing with very high-dimensional Bermudan-style derivatives. For such problems, we quickly see that we can at best hope for price bounds, and we can only use a simulation approach.…
We study a speculative trading problem within the exploratory reinforcement learning (RL) framework of Wang et al. [2020]. The problem is formulated as a sequential optimal stopping problem over entry and exit times under general utility…
Sample-efficient exploration is crucial not only for discovering rewarding experiences but also for adapting to environment changes in a task-agnostic fashion. A principled treatment of the problem of optimal input synthesis for system…
We develop a continuous-time reinforcement learning framework for a class of singular stochastic control problems without entropy regularization. The optimal singular control is characterized as the optimal singular control law, which is a…
In this paper we study by probabilistic techniques the convergence of the value function for a two-scale, infinite-dimensional, stochastic controlled system as the ratio between the two evolution speeds diverges. The value function is…
We study the exploration problem with approximate linear action-value functions in episodic reinforcement learning under the notion of low inherent Bellman error, a condition normally employed to show convergence of approximate value…
We study reflected backward stochastic differential equation (RBSDEs) on the probability space equipped with a Brownian motion. The main novelty of the paper lies in fact that we consider the following weak assumptions on the data: barriers…
We study an optimal control problem on infinite horizon for a controlled stochastic differential equation driven by Brownian motion, with a discounted reward functional. The equation may have memory or delay effects in the coefficients,…
This paper investigates an optimal consumption-investment problem featuring recursive utility via Tsallis relative entropy. We establish a fundamental connection between this optimization problem and a quadratic backward stochastic…
In this paper we first investigate zero-sum two-player stochastic differential games with reflection with the help of theory of Reflected Backward Stochastic Differential Equations (RBSDEs). We will establish the dynamic programming…
This paper proposes a relaxed control regularization with general exploration rewards to design robust feedback controls for multi-dimensional continuous-time stochastic exit time problems. We establish that the regularized control problem…
We study mean-field games of optimal stopping (OS-MFGs) and introduce an entropy-regularized framework to enable learning-based solution methods. By utilizing randomized stopping times, we reformulate the OS-MFG as a mean-field game of…
We consider the problem of learning the optimal policy for Markov decision processes with safety constraints. We formulate the problem in a reach-avoid setup. Our goal is to design online reinforcement learning algorithms that ensure safety…
In this paper, we introduce a non-linear Snell envelope which at each time represents the maximal value that can be achieved by stopping a BSDE with constrained jumps. We establish the existence of the Snell envelope by employing a…
Considering uncertainties and disturbances is an important, yet challenging, step in successful decision making. The problem becomes more challenging in safety-constrained environments. In this paper, we propose a robust and safe trajectory…
We consider a classical finite horizon optimal control problem for continuous-time pure jump Markov processes described by means of a rate transition measure depending on a control parameter and controlled by a feedback law. For this class…
We investigate an optimal stopping problem for the expected value of a discounted payoff on a regime-switching geometric Brownian motion under two constraints on the possible stopping times: only at exogenous random times and only during a…
We construct an aggregated version of the value processes associated with stochastic control problems, where the criterion to optimise is given by solutions to semi-martingale backward stochastic differential equations (BSDEs). The results…