Related papers: Actor-Critic Algorithm for High-dimensional Partia…
We introduce a reinforcement learning method for a class of non-Markov systems; our approach extends the actor-critic framework given by Rose et al. [New J. Phys. 23 013013 (2021)] for obtaining scaled cumulant generating functions…
We extend the Deep Galerkin Method (DGM) introduced in Sirignano and Spiliopoulos (2018)} to solve a number of partial differential equations (PDEs) that arise in the context of optimal stochastic control and mean field games. First, we…
We present the first class of policy-gradient algorithms that work with both state-value and policy function-approximation, and are guaranteed to converge under off-policy training. Our solution targets problems in reinforcement learning…
We develop a parameterized Primal-Dual $\pi$ Learning method based on deep neural networks for Markov decision process with large state space and off-policy reinforcement learning. In contrast to the popular Q-learning and actor-critic…
Partial Differential Equations (PDEs) are used to model a variety of dynamical systems in science and engineering. Recent advances in deep learning have enabled us to solve them in a higher dimension by addressing the curse of…
We are interested in stochastic control problems coming from mathematical finance and, in particular, related to model uncertainty, where the uncertainty affects both volatility and intensity. This kind of stochastic control problems is…
We propose algorithms for solving high-dimensional Partial Differential Equations (PDEs) that combine a probabilistic interpretation of PDEs, through Feynman-Kac representation, with sparse interpolation. Monte-Carlo methods and…
Stemmed from the derivation of the optimal control to a stochastic linear-quadratic control problem with Markov jumps, we study one kind of backward stochastic differential equations (BSDEs) that the generator f is affected by a Markovian…
Fractional Brownian motions(fBMs) are not semimartingales so the classical theory of It\^o integral can't apply to fBMs. Wick integration as one of the applications of Malliavin calculus to stochastic analysis is a fine definition for fBMs.…
Multi-agent deep reinforcement learning has been applied to address a variety of complex problems with either discrete or continuous action spaces and achieved great success. However, most real-world environments cannot be described by only…
In this work we study the numerical approximation of a class of ergodic Backward Stochastic Differential Equations. These equations are formulated in an infinite horizon framework and provide a probabilistic representation for elliptic…
We present a non-asymptotic convergence analysis of $Q$-learning and actor-critic algorithms for robust average-reward Markov Decision Processes (MDPs) under contamination, total-variation (TV) distance, and Wasserstein uncertainty sets. A…
Policy gradient methods in actor-critic reinforcement learning (RL) have become perhaps the most promising approaches to solving continuous optimal control problems. However, the trial-and-error nature of RL and the inherent randomness…
We study a class of backward doubly stochastic differential equations (BDSDEs) involving martingales with spatial parameters, and show that they provide probabilistic interpretations (Feynman-Kac formulae) for certain semilinear stochastic…
In recent years, tremendous progress has been made on numerical algorithms for solving partial differential equations (PDEs) in a very high dimension, using ideas from either nonlinear (multilevel) Monte Carlo or deep learning. They are…
Recently, the deep learning method has been used for solving forward-backward stochastic differential equations (FBSDEs) and parabolic partial differential equations (PDEs). It has good accuracy and performance for high-dimensional…
It is known that Markovian forward-backward stochastic differential equations provide nonlinear Feynman-Kac representation formulae for semilinear parabolic PDEs. We show that non-Markovian forward-backward stochastic differential equations…
Despite definite success in deep reinforcement learning problems, actor-critic algorithms are still confronted with sample inefficiency in complex environments, particularly in tasks where efficient exploration is a bottleneck. These…
To learn approximately optimal acting policies for decision problems, modern Actor Critic algorithms rely on deep Neural Networks (DNNs) to parameterize the acting policy and greedification operators to iteratively improve it. The reliance…
Parabolic partial differential equations (PDEs) and backward stochastic differential equations (BSDEs) have a wide range of applications. In particular, high-dimensional PDEs with gradient-dependent nonlinearities appear often in the…