Related papers: An Entropy Regularized BSDE Approach to Bermudan O…
In this paper we study, by probabilistic techniques, the convergence of the value function for a two-scale, infinite-dimensional, stochastic controlled system as the ratio between the two evolution speeds diverges. The value function is…
We study penalization coupled with time discretization for decoupled Markovian doubly reflected BSDEs with obstacles \(p_b(t,X_t)\le Y_t\le p_w(t,X_t)\). The DRBSDE is approximated by a penalized BSDE with parameter \(\lambda\) and…
The problem of pricing Bermudan options using Monte Carlo and a nonparametric regression is considered. We derive optimal non-asymptotic bounds for a lower biased estimate based on the suboptimal stopping rule constructed using some…
We approach the continuous-time mean-variance (MV) portfolio selection with reinforcement learning (RL). The problem is to achieve the best tradeoff between exploration and exploitation, and is formulated as an entropy-regularized, relaxed…
Sample-based trajectory optimisers are a promising tool for the control of robotics with non-differentiable dynamics and cost functions. Contemporary approaches derive from a restricted subclass of stochastic optimal control where the…
We propose a general framework for entropy-regularized average-reward reinforcement learning in Markov decision processes (MDPs). Our approach is based on extending the linear-programming formulation of policy optimization in MDPs to…
Efficient exploration is a central problem in reinforcement learning and is often formalized as maximizing the entropy of the state-action occupancy measure. While unconstrained maximum-entropy exploration is relatively well understood,…
We address the challenge of exploration in reinforcement learning (RL) when the agent operates in an unknown environment with sparse or no rewards. In this work, we study the maximum entropy exploration problem of two different types. The…
We consider reflected backward stochastic different equations with optional barrier and so-called regulated trajectories, i.e trajectories with left and right finite limits. We prove existence and uniqueness results. We also show that the…
We consider the optimal stopping problem with non-linear $f$-expectation (induced by a BSDE) without making any regularity assumptions on the reward process $\xi$. and with general filtration. We show that the value family can be aggregated…
In this paper, we study the optimal dividend problem under the continuous time diffusion model with the bounded dividend rate from the Reinforcement Learning (RL) perspective. Unlike the standard literature, our main focus will be on…
Two hitherto disconnected threads of research, diverse exploration (DE) and maximum entropy RL have addressed a wide range of problems facing reinforcement learning algorithms via ostensibly distinct mechanisms. In this work, we identify a…
Efficient exploration remains a central challenge in reinforcement learning, serving as a useful pretraining objective for data collection, particularly when an external reward function is unavailable. A principled formulation of the…
This paper studies finite-time optimal consumption-investment problems with power, logarithmic and exponential utilities, in a regime switching market with random coefficients, subject to coupled constraints on the consumption and…
The present paper studies a kind of robust optimization problems with constraint. The problem is formulated through Backward Stochastic Differential Equations (BSDEs) with quadratic generators. A necessary condition is established for the…
We study the properties of the free boundaries and the corresponding hitting times in the context of optimal stopping in discrete time. We first prove the continuity of the map from the boundaries to the expected value of the corresponding…
We present a novel method for solving a class of time-inconsistent optimal stopping problems by reducing them to a family of standard stochastic optimal control problems. In particular, we convert an optimal stopping problem with a…
This work uses the entropy-regularised relaxed stochastic control perspective as a principled framework for designing reinforcement learning (RL) algorithms. Herein agent interacts with the environment by generating noisy controls…
We introduce a new type of reflected backward stochastic differential equations (BSDEs) for which the reflection constraint is imposed on its main solution component, denoted as $Y$ by convention, but in terms of its conditional expectation…
We study sequential decision-making when the agent's internal model class is misspecified. Within the infinite-horizon Berk-Nash framework, stable behavior arises as a fixed point: the agent acts optimally relative to a subjective model,…