Related papers: Is RL fine-tuning harder than regression? A PDE le…
This work is devoted to the study of optimal control of stochastic functional differential equations (SFDEs) and its application to mathematical finance. By using the Dynkin formula and solution of the Dirichlet-Poisson problem, the…
Diffusion models recently emerged as a powerful paradigm for recommender systems, offering state-of-the-art performance by modeling the generative process of user-item interactions. However, training such models from scratch is both…
We treat infinite horizon optimal control problems by solving the associated stationary Hamilton-Jacobi-Bellman (HJB) equation numerically to compute the value function and an optimal feedback law. The dynamical systems under consideration…
Fine-tuning pretrained language models (PLMs) on downstream tasks has become common practice in natural language processing. However, most of the PLMs are vulnerable, e.g., they are brittle under adversarial attacks or imbalanced data,…
Policy gradient (PG) methods are successful approaches to deal with continuous reinforcement learning (RL) problems. They learn stochastic parametric (hyper)policies by either exploring in the space of actions or in the space of parameters.…
In performative Reinforcement Learning (RL), an agent faces a policy-dependent environment: the reward and transition functions depend on the agent's policy. Prior work on performative RL has studied the convergence of repeated retraining…
We consider the problem of learning models for risk-sensitive reinforcement learning. We theoretically demonstrate that proper value equivalence, a method of learning models which can be used to plan optimally in the risk-neutral setting,…
We investigate the long time behavior of weakly dissipative semilinear Hamilton-Jacobi-Bellman (HJB) equations and the turnpike property for the corresponding stochastic control problems. To this aim, we develop a probabilistic approach…
We consider the numerical solution of Hamilton-Jacobi-Bellman equations arising in stochastic control theory. We introduce a class of monotone approximation schemes relying on monotone interpolation. These schemes converge under very weak…
The focus of this article is studying an optimal control problem for branching diffusion processes. Initially, we introduce the problem in its strong formulation and expand it to include linearly growing drifts. Then, we present a relaxed…
We explore the approximation of feedback control of integro-differential equations containing a fractional Laplacian term. To obtain feedback control for the state variable of this nonlocal equation we use the Hamilton--Jacobi--Bellman…
This paper studies an optimal dividend problem with a drawdown constraint in a Brownian motion model, requiring the dividend payout rate to remain above a fixed proportion of its historical maximum. This leads to a path-dependent stochastic…
Convex Q-learning is a recent approach to reinforcement learning, motivated by the possibility of a firmer theory for convergence, and the possibility of making use of greater a priori knowledge regarding policy or value function structure.…
We study a stochastic optimal control problem for a partially observed diffusion. By using the control randomization method in [4], we prove a corresponding randomized dynamic programming principle (DPP) for the value function, which is…
It has been recognized that the data generated by the denoising diffusion probabilistic model (DDPM) improves adversarial training. After two years of rapid development in diffusion models, a question naturally arises: can better diffusion…
Robustness to modeling errors and uncertainties remains a central challenge in reinforcement learning (RL). In this work, we address this challenge by leveraging diffusion models to train robust RL policies. Diffusion models have recently…
Model-based Reinforcement Learning estimates the true environment through a world model in order to approximate the optimal policy. This family of algorithms usually benefits from better sample efficiency than their model-free counterparts.…
We study off-policy reinforcement learning for controlling continuous-time Markov diffusion processes with discrete-time observations and actions. We consider model-free algorithms with function approximation that learn value and advantage…
In this work, we propose a class of numerical schemes for solving semilinear Hamilton-Jacobi-Bellman-Isaacs (HJBI) boundary value problems which arise naturally from exit time problems of diffusion processes with controlled drift. We…
In this paper, we introduce Hamilton-Jacobi-Bellman (HJB) equations for Q-functions in continuous time optimal control problems with Lipschitz continuous controls. The standard Q-function used in reinforcement learning is shown to be the…