Related papers: An Entropy Regularized BSDE Approach to Bermudan O…
Reinforcement learning (RL) for exponential-utility optimization in discounted Markov decision processes (MDPs) lacks principled value-based algorithms. We address this gap in the fixed risk-aversion setting. Building on the Bellman-type…
Nowadays many financial derivatives, such as American or Bermudan options, are of early exercise type. Often the pricing of early exercise options gives rise to high-dimensional optimal stopping problems, since the dimension corresponds to…
In this paper, we focus on the problem of robustifying reinforcement learning (RL) algorithms with respect to model uncertainties. Indeed, in the framework of model-based RL, we propose to merge the theory of constrained Markov decision…
In this paper, an optimal switching problem is proposed for one-dimensional reflected backward stochastic differential equations (RBSDEs, for short) where the generators, the terminal values and the barriers are all switched with positive…
In this paper we consider a method of solving optimal stopping problems in discrete and continuous time based on their dual representation. A novel and generic simulation-based optimization algorithm not involving nested simulations is…
In the first part of this paper, we study RBSDEs in the case where the filtration is not quasi-left continuous and the lower obstacle is given by a predictable process. We prove the existence and uniqueness by using some results of optimal…
We study an optimal execution problem in illiquid markets with both instantaneous and persistent price impact and stochastic resilience when only absolutely continuous trading strategies are admissible. In our model the value function can…
This paper shows that penalized backward stochastic differential equation (BSDE), which is often used to approximate and solve the corresponding reflected BSDE, admits both optimal stopping representation and optimal control representation.…
Reinforcement Learning (RL) has emerged as a powerful framework for sequential decision-making in dynamic environments, particularly when system parameters are unknown. This paper investigates RL-based control for entropy-regularized…
Reinforcement learning (RL) is a powerful framework for decision-making in uncertain environments, but it often requires large amounts of data to learn an optimal policy. We address this challenge by incorporating prior model knowledge to…
We study a class of backward stochastic differential equations (BSDEs) driven by a random measure or, equivalently, by a marked point process. Under appropriate assumptions we prove well-posedness and continuous dependence of the solution…
In this paper we introduce a numerical method for optimal stopping in the framework of one dimensional diffusion. We use the Skorokhod embedding in order to construct recombining tree approximations for diffusions with general coefficients.…
We first study an optimal stopping problem in which a player (an agent) uses a discrete stopping time in order to stop optimally a payoff process whose risk is evaluated by a (non-linear) $g$-expectation. We then consider a non-zero-sum…
In this paper we make a survey on the so called randomization method, a recent methodology to study stochastic optimization problems. It allows to represent the value function of an optimal control problem by a suitable backward stochastic…
One of the most critical challenges in deep reinforcement learning is to maintain the long-term exploration capability of the agent. To tackle this problem, it has been recently proposed to provide intrinsic rewards for the agent to…
In this paper, we study reflected backward stochastic difference equations (RBSDEs for short) with finitely many states in discrete time. The general existence and uniqueness result, as well as comparison theorems for the solutions, are…
We consider a robust impulse control problem in finite horizon where the underlying uncertainty stems from an impulsively and continuously controlled functional stochastic differential equation (FSDE) driven by Brownian motion. We assume…
We investigate an optimal prevention and insurance problem in a general risk setting, where a representative agent is exposed to potential losses. The agent adopts a strategy that combines self-protection, aimed at reducing the frequency of…
In online reinforcement learning, data scarcity creates epistemic uncertainty that makes robustness important early in learning, whereas sufficient exploration is needed to learn the true-environment optimal policy. We study this…
In this paper, it is shown that Bermudan option pricing based on either the r\'eduite (in a one-dimensional setting: piecewise harmonic interpolation) or cubature -- is sensible from an economic vantage point: Any sequence of thus-computed…