English
Related papers

Related papers: An Entropy Regularized BSDE Approach to Bermudan O…

200 papers

Safe exploration is crucial for the real-world application of reinforcement learning (RL). Previous works consider the safe exploration problem as Constrained Markov Decision Process (CMDP), where the policies are being optimized under…

Machine Learning · Computer Science 2021-07-12 Hao Sun , Ziping Xu , Meng Fang , Zhenghao Peng , Jiadong Guo , Bo Dai , Bolei Zhou

In this paper, we consider dynamic risk measures induced by backward stochastic differential equations (BSDEs). We discuss different examples that come up in the literature, including the entropic risk measure and the risk measure arising…

Probability · Mathematics 2024-08-07 Nacira Agram , Jan Rems , Emanuela Rosazza Gianin

We study risk-sensitive reinforcement learning in finite discounted MDPs with recursive entropic risk measures (ERM), where the risk parameter $\beta \neq 0$ controls the agent's risk attitude: $\beta>0$ for risk-averse and $\beta<0$ for…

Machine Learning · Computer Science 2026-05-20 Oliver Mortensen , Mohammad Sadegh Talebi

Inference scaling helps LLMs solve complex reasoning problems through extended runtime computation. On top of long chain-of-thought (long-CoT) models, purely inference-time techniques such as best-of-N (BoN) sampling, majority voting, or…

We present a random measure approach for modeling exploration, i.e., the execution of measure-valued controls, in continuous-time reinforcement learning (RL) with controlled diffusion and jumps. First, we consider the case when sampling the…

Machine Learning · Computer Science 2024-09-27 Christian Bender , Nguyen Tran Thuan

This paper presents a new formulation for model-free robust optimal regulation of continuous-time nonlinear systems. The proposed reinforcement learning based approach, referred to as incremental adaptive dynamic programming (IADP),…

Systems and Control · Electrical Eng. & Systems 2022-03-25 Cong Li , Yongchao Wang , Fangzhou Liu , Qingchen Liu , Martin Buss

A promising technique for exploration is to maximize the entropy of visited state distribution, i.e., state entropy, by encouraging uniform coverage of visited state space. While it has been effective for an unsupervised setup, it tends to…

Machine Learning · Computer Science 2024-08-12 Dongyoung Kim , Jinwoo Shin , Pieter Abbeel , Younggyo Seo

We propose a new approach to solve optimal stopping problems via simulation. Working within the backward dynamic programming/Snell envelope framework, we augment the methodology of Longstaff-Schwartz that focuses on approximating the…

Computational Finance · Quantitative Finance 2015-09-04 Robert B. Gramacy , Mike Ludkovski

This paper presents a novel and direct approach to price boundary and final-value problems, corresponding to barrier options, using forward deep learning to solve forward-backward stochastic differential equations (FBSDEs). Barrier…

Computational Finance · Quantitative Finance 2024-09-13 Narayan Ganesan , Yajie Yu , Bernhard Hientzsch

Regularization of control policies using entropy can be instrumental in adjusting predictability of real-world systems. Applications benefiting from such approaches range from, e.g., cybersecurity, which aims at maximal unpredictability, to…

Systems and Control · Electrical Eng. & Systems 2026-02-18 Menno van Zutphen , Giannis Delimpaltadakis , Maurice Heemels , Duarte Antunes

Entropic regularization of policies in Reinforcement Learning (RL) is a commonly used heuristic to ensure that the learned policy explores the state-space sufficiently before overfitting to a local optimal policy. The primary motivation for…

Machine Learning · Computer Science 2021-01-19 Hisham Husain , Kamil Ciosek , Ryota Tomioka

We consider a class of time-inhomogeneous optimal stopping problems and we provide sufficient conditions on the data of the problem that guarantee monotonicity of the optimal stopping boundary. In our setting, time-inhomogeneity stems not…

Optimization and Control · Mathematics 2023-01-16 Alessandro Milazzo

We present a new deep primal-dual backward stochastic differential equation framework based on stopping time iteration to solve optimal stopping problems. A novel loss function is proposed to learn the conditional expectation, which…

Computational Finance · Quantitative Finance 2024-09-12 Jiefei Yang , Guanglian Li

We study the stochastic control-stopping problem when the data are of polynomial growth. The approach is based on backward stochastic dierential equations (BSDEs for short). The problem turns into the study of a specic reected BSDE with a…

Optimization and Control · Mathematics 2020-05-15 Brahim Asri , Said Hamadène , Khalid Oufdil

We study the optimal liquidation problems in target zone models using dynamic programming methods. Such control problems allow for stochastic differential equations with reflections and random coefficients. The value function is…

Optimization and Control · Mathematics 2019-12-17 Robert Elliott , Jinniao Qiu , Wenning Wei

Reinforcement Learning algorithms are primarily focused on learning a policy that maximizes expected return. As a result, the learned policy can exploit one or few reward sources. However, in many natural situations, it is desirable to…

Machine Learning · Computer Science 2026-03-31 Sagalpreet Singh , Rishi Saket , Aravindan Raghuveer

In this paper we study simulation based optimization algorithms for solving discrete time optimal stopping problems. This type of algorithms became popular among practioneers working in the area of quantitative finance. Using large…

Optimization and Control · Mathematics 2009-09-22 Denis Belomestny

We introduce and study a new class of optimal switching problems, namely switching problem with controlled randomisation, where some extra-randomness impacts the choice of switching modes and associated costs. We show that the optimal value…

Probability · Mathematics 2020-01-31 Cyril Bénézet , Jean-François Chassagneux , Adrien Richou

In deterministic systems, reinforcement learning-based online approximate optimal control methods typically require a restrictive persistence of excitation (PE) condition for convergence. This paper presents a concurrent learning-based…

Systems and Control · Computer Science 2017-07-25 Rushikesh Kamalapurkar , Patrick Walters , Warren Dixon

Sequential Bayesian optimal experimental design (SBOED) for PDE-governed inverse problems is computationally challenging, especially for infinite-dimensional random field parameters. High-fidelity approaches require repeated forward and…

Optimization and Control · Mathematics 2026-01-12 Kaichen Shen , Peng Chen