English
Related papers

Related papers: An Entropy Regularized BSDE Approach to Bermudan O…

200 papers

We investigate two new strategies for the numerical solution of optimal stopping problems within the Regression Monte Carlo (RMC) framework of Longstaff and Schwartz. First, we propose the use of stochastic kriging (Gaussian process)…

Computational Finance · Quantitative Finance 2016-10-27 Michael Ludkovski

We address the problem of computing reliable policies in reinforcement learning problems with limited data. In particular, we compute policies that achieve good returns with high confidence when deployed. This objective, known as the…

Machine Learning · Computer Science 2021-03-01 Bahram Behzadian , Reazul Hasan Russel , Marek Petrik , Chin Pang Ho

We consider dynamic risk measures induced by Backward Stochastic Differential Equations (BSDEs) in enlargement of filtration setting. On a fixed probability space, we are given a standard Brownian motion and a pair of random variables…

Risk Management · Quantitative Finance 2020-09-25 Alessandro Calvia , Emanuela Rosazza Gianin

Resolving the exploration-exploitation trade-off remains a fundamental problem in the design and implementation of reinforcement learning (RL) algorithms. In this paper, we focus on model-free RL using the epsilon-greedy exploration policy,…

Machine Learning · Computer Science 2020-07-03 Michael Gimelfarb , Scott Sanner , Chi-Guhn Lee

In practice, the parameters of control policies are often tuned manually. This is time-consuming and frustrating. Reinforcement learning is a promising alternative that aims to automate this process, yet often requires too many experiments…

We introduce Random Reward Perturbation (RRP), a novel exploration strategy for reinforcement learning (RL). Our theoretical analyses demonstrate that adding zero-mean noise to environmental rewards effectively enhances policy diversity…

Machine Learning · Computer Science 2025-06-11 Haozhe Ma , Guoji Fu , Zhengding Luo , Jiele Wu , Tze-Yun Leong

We propose Deterministic Sequencing of Exploration and Exploitation (DSEE) algorithm with interleaving exploration and exploitation epochs for model-based RL problems that aim to simultaneously learn the system model, i.e., a Markov…

Machine Learning · Computer Science 2022-12-21 Piyush Gupta , Vaibhav Srivastava

With the growing needs of online A/B testing to support the innovation in industry, the opportunity cost of running an experiment becomes non-negligible. Therefore, there is an increasing demand for an efficient continuous monitoring…

Machine Learning · Computer Science 2023-04-04 Runzhe Wan , Yu Liu , James McQueen , Doug Hains , Rui Song

With the increasing pace of automation, modern robotic systems need to act in stochastic, non-stationary, partially observable environments. A range of algorithms for finding parameterized policies that optimize for long-term average…

Machine Learning · Computer Science 2019-09-04 David Nass , Boris Belousov , Jan Peters

In this paper, we consider a modified version of the control problem in a model free Markov decision process (MDP) setting with large state and action spaces. The control problem most commonly addressed in the contemporary literature is to…

Artificial Intelligence · Computer Science 2018-02-01 Ajin George Joseph , Shalabh Bhatnagar

In this article we continue our investigation of the iterative regularization method for optimization problems based on Bregman distances. The optimization problems are subject to pointwise inequality constraints in $L^2(\Omega)$. We…

Optimization and Control · Mathematics 2016-08-25 Frank Pörner

This paper addresses the existence and uniqueness of solutions to Reflected Generalized Backward Stochastic Differential Equations (GRBSDEs) within a general filtration that supports a Brownian motion and an independent integer-valued…

Probability · Mathematics 2026-03-09 Badr Elmansouri , Mohamed El Otmani

We study a robust utility maximization problem in the unbounded case with a general penalty term and information including jumps. We focus on time consistent penalties and we prove that there exists an optimal probability measure solution…

Optimization and Control · Mathematics 2022-12-07 Sarah Kaakai , Anis Matoussi , Achraf Tamtalini

This paper is dedicated to the analysis of backward stochastic differential equations (BSDEs) with jumps, subject to an additional global constraint involving all the components of the solution. We study the existence and uniqueness of a…

Probability · Mathematics 2011-03-10 Romuald Elie , Idris Kharroubi

Practical reinforcement learning problems are often formulated as constrained Markov decision process (CMDP) problems, in which the agent has to maximize the expected return while satisfying a set of prescribed safety constraints. In this…

Machine Learning · Computer Science 2019-09-23 Shin-ichi Maeda , Hayato Watahiki , Shintarou Okada , Masanori Koyama

This paper establishes a rigorous connection between regularized discrete-time reinforcement learning (RL) and continuous-time stochastic optimal control. Specifically, classical RL algorithms are typically solving a regularized…

Optimization and Control · Mathematics 2026-04-24 Huyên Pham , Yuming Paul Zhang , Yuhua Zhu

Benders decomposition (BD), along with its generalized version (GBD), is a widely used algorithm for solving large-scale mixed-integer optimization problems that arise in the operation of process systems. However, the off-the-shelf…

Optimization and Control · Mathematics 2025-08-12 Zhe Li , Bernard T. Agyeman , Ilias Mitrai , Prodromos Daoutidis

In this paper, we investigate the optimal output tracking problem for linear discrete-time systems with unknown dynamics using reinforcement learning and robust output regulation theory. This output tracking problem only allows to utilize…

Dynamical Systems · Mathematics 2021-01-22 Ci Chen , Lihua Xie , Yi Jiang , Kan Xie , Shengli Xie

In this paper we study different algorithms for reflected backward stochastic differential equations (BSDE in short) with two continuous barriers basing on random work framework. We introduce different numerical algorithms by penalization…

Probability · Mathematics 2009-09-23 Mingyu Xu

We consider reflected backward stochastic differential equations with two optional barriers of class (D) satisfying Mokobodzki's separation condition and coefficient which is only continuous and non-increasing. We assume that data are…

Probability · Mathematics 2021-12-02 Tomasz Klimsiak , Maurycy Rzymowski