English
Related papers

Related papers: An Entropy Regularized BSDE Approach to Bermudan O…

200 papers

In this paper, we study policy evaluation in continuous-time reinforcement learning (RL), where the state follows an unknown stochastic differential equation (SDE), but only discrete-time data are available. We first highlight that the…

Optimization and Control · Mathematics 2026-02-23 Yuhua Zhu

We study multi-objective reinforcement learning with nonlinear preferences over trajectories. That is, we maximize the expected value of a nonlinear function over accumulated rewards (expected scalarized return or ESR) in a multi-objective…

Machine Learning · Computer Science 2025-02-19 Nianli Peng , Muhang Tian , Brandon Fain

This paper focuses on stochastic optimal control problems with constraints in law, which are rewritten as optimization (minimization) of probability measures problem on the canonical space. We introduce a penalized version of this type of…

Optimization and Control · Mathematics 2025-03-18 Thibaut Bourdais , Nadia Oudjane , Francesco Russo

This paper is concerned with a stochastic recursive optimal control problem with time delay, where the controlled system is described by a stochastic differential delayed equation (SDDE) and the cost functional is formulated as the solution…

Optimization and Control · Mathematics 2014-08-26 Jingtao Shi , Huanshui Zhang

We study a class of reflected backward stochastic differential equations with nonpositive jumps and upper barrier. Existence and uniqueness of a minimal solution is proved by a double penalization approach under regularity assumptions on…

Probability · Mathematics 2013-08-27 Sébastien Choukroun , Andrea Cosso , Huyen Pham

Reinforcement learning is widely used in applications where one needs to perform sequential decisions while interacting with the environment. The problem becomes more challenging when the decision requirement includes satisfying some safety…

Machine Learning · Computer Science 2022-07-15 Qinbo Bai , Amrit Singh Bedi , Mridul Agarwal , Alec Koppel , Vaneet Aggarwal

We study a finite-inventory risk-sensitive market making problem in which a dealer controls bid and ask quotes, faces Brownian midprice risk, and receives liquidity-taking orders through point processes with quote-dependent intensities. The…

Trading and Market Microstructure · Quantitative Finance 2026-05-26 Tenghan Zhong

Robust Markov decision processes (RMDPs) extend standard Markov decision processes (MDPs) to account for uncertainty in the transition probabilities. RMDPs have an uncertainty set that defines a set of possible transition functions, each of…

Logic in Computer Science · Computer Science 2026-04-30 Marnix Suilen , Guillermo A. Pérez

Reinforcement learning algorithms typically consider discrete-time dynamics, even though the underlying systems are often continuous in time. In this paper, we introduce a model-based reinforcement learning algorithm that represents…

Machine Learning · Computer Science 2023-11-01 Lenart Treven , Jonas Hübotter , Bhavya Sukhija , Florian Dörfler , Andreas Krause

We study a discrete time approximation scheme for the solution of a doubly reflected Backward Stochastic Differential Equation (DBBSDE in short) with jumps, driven by a Brownian motion and an independent compensated Poisson process.…

Probability · Mathematics 2016-12-14 Roxana Dumitrescu , Céline Labart

Within the framework of probably approximately correct Markov decision processes (PAC-MDP), much theoretical work has focused on methods to attain near optimality after a relatively long period of learning and exploration. However,…

Artificial Intelligence · Computer Science 2016-04-06 Kenji Kawaguchi

We introduce a framework for approximate dynamic programming that we apply to discrete time chains on $\mathbb{Z}_+^d$ with countable action sets. Our approach is grounded in the approximation of the (controlled) chain's generator by that…

Optimization and Control · Mathematics 2018-04-16 Anton Braverman , Itai Gurvich , Junfei Huang

Despite the close connection between exploration and sample efficiency, most state of the art reinforcement learning algorithms include no considerations for exploration beyond maximizing the entropy of the policy. In this work we address…

This paper studies the mean-field Markov decision process (MDP) with the centralized stopping under the non-exponential discount. The problem differs fundamentally from most existing studies on mean-field optimal control/stopping due to its…

Optimization and Control · Mathematics 2025-01-22 Xiang Yu , Fengyi Yuan

This paper studies stochastic control problems motivated by optimal consumption with wealth benchmark tracking. The benchmark process is modeled by a combination of a geometric Brownian motion and a running maximum process, indicating its…

Optimization and Control · Mathematics 2024-04-26 Lijun Bo , Yijie Huang , Xiang Yu

Model-free deep-reinforcement-based learning algorithms have been applied to a range of COPs~\cite{bello2016neural}~\cite{kool2018attention}~\cite{nazari2018reinforcement}. However, these approaches suffer from two key challenges when…

Machine Learning · Computer Science 2022-06-01 Nasrin Sultana , Jeffrey Chan , Tabinda Sarwar , A. K. Qin

In this paper, we provide two new stable online algorithms for the problem of prediction in reinforcement learning, \emph{i.e.}, estimating the value function of a model-free Markov reward process using the linear function approximation…

Machine Learning · Computer Science 2018-06-19 Ajin George Joseph , Shalabh Bhatnagar

We investigate optimal stopping problems for systems driven by the Brownian sheet. Our analysis is divided into two parts. In the first part we derive explicit solutions to two optimal stopping problems for the exponentially discounted…

Probability · Mathematics 2026-03-16 Nacira Agram , Bernt Oksendal , Frank Proske , Olena Tymoshenko

We study backward stochastic differential equations (BSDEs) for time-changed L\'evy noises when the time-change is independent of the L\'evy process. We prove existence and uniqueness of the solution and we obtain an explicit formula for…

Probability · Mathematics 2013-12-19 Giulia Di Nunno , Steffen Sjursen

Model-based reinforcement learning (MBRL) is believed to have higher sample efficiency compared with model-free reinforcement learning (MFRL). However, MBRL is plagued by dynamics bottleneck dilemma. Dynamics bottleneck dilemma is the…

Machine Learning · Computer Science 2021-06-25 Xiyao Wang , Junge Zhang , Wenzhen Huang , Qiyue Yin
‹ Prev 1 8 9 10 Next ›