Related papers: Entropy-regularized penalization schemes and refle…
We investigate the existence and uniqueness of (locally) absolutely continuous trajectories of a penalty term-based dynamical system associated to a constrained variational inequality expressed as a monotone inclusion problem. Relying on…
This paper develops the first class of algorithms that enable unbiased estimation of steady-state expectations for multidimensional reflected Brownian motion. In order to explain our ideas, we first consider the case of compound Poisson…
When deploying artificial agents in real-world environments where they interact with humans, it is crucial that their behavior is aligned with the values, social norms or other requirements of that environment. However, many environments…
Many offline unsupervised change point detection algorithms rely on minimizing a penalized sum of segment-wise costs. We extend this framework by proposing to minimize a sum of discrepancies between segments. In particular, we propose to…
We address the challenge of exploration in reinforcement learning (RL) when the agent operates in an unknown environment with sparse or no rewards. In this work, we study the maximum entropy exploration problem of two different types. The…
We propose a new methodology for parameterized constrained robust optimization, an important class of optimization problems under uncertainty, based on learning with a self-supervised penalty-based loss function. Whereas supervised learning…
Differential Dynamic Programming (DDP) has become a well established method for unconstrained trajectory optimization. Despite its several applications in robotics and controls however, a widely successful constrained version of the…
In this paper, we study the backward stochastic differential equation (BSDE) with two nonlinear mean reflections, which means that the constraints are imposed on the distribution of the solution but not on its paths. Based on the backward…
The present paper is devoted to the study of backward stochastic differential equations with mean reflection formulated by Briand et al. [7]. We investigate the solvability of a generalized mean reflected BSDE, whose driver also depends on…
In this paper, we study the reflected BSDE with one continuous barrier, under the monotonicity and general increasing condition on $y$ and non Lipschitz condition on $z$. We prove the existence and uniqueness of the solution to these…
We develop a continuous-time reinforcement learning framework for a class of singular stochastic control problems without entropy regularization. The optimal singular control is characterized as the optimal singular control law, which is a…
The problem of optimal stopping with finite horizon in discrete time is considered in view of maximizing the expected gain. The algorithm proposed in this paper is completely nonparametric in the sense that it uses observed data from the…
Reinforcement learning (RL) is an important field of research in machine learning that is increasingly being applied to complex optimization problems in physics. In parallel, concepts from physics have contributed to important advances in…
We propose a new online algorithm for cumulative regret minimization in a stochastic linear bandit. The algorithm pulls the arm with the highest estimated reward in a linear model trained on its perturbed history. Therefore, we call it…
Sample-efficient exploration is crucial not only for discovering rewarding experiences but also for adapting to environment changes in a task-agnostic fashion. A principled treatment of the problem of optimal input synthesis for system…
We study the well-posedness of general reflected BSDEs driven by a continuous martingale, when the coefficient f of the driver has at most quadratic growth in the control variable Z, with a bounded terminal condition and a lower obstacle…
Maximum entropy reinforcement learning integrates exploration into policy learning by providing additional intrinsic rewards proportional to the entropy of some distribution. In this paper, we propose a novel approach in which the intrinsic…
Many inverse and parameter estimation problems can be written as PDE-constrained optimization problems. The goal, then, is to infer the parameters, typically coefficients of the PDE, from partial measurements of the solutions of the PDE for…
Sparse estimation methods are aimed at using or obtaining parsimonious representations of data or models. They were first dedicated to linear variable selection but numerous extensions have now emerged such as structured sparsity or kernel…
We solve a class of doubly reflected backward stochastic differential equation whose generator depends on the resistance due to reflections, which extend the recent work of Qian and Xu on reflected BSDE with one barrier. We then obtain the…