中文
相关论文

相关论文: Actor-Critic Algorithm for High-dimensional Partia…

200 篇论文

In this work we propose a new algorithm for solving high-dimensional backward stochastic differential equations (BSDEs). Based on the general theta-discretization for the time-integrands, we show how to efficiently use eXtreme Gradient…

数值分析 · 数学 2021-07-15 Long Teng

Forward-backward stochastic differential equations (FBSDEs) have attracted significant attention since they were introduced almost 30 years ago, due to their wide range of applications, from solving non-linear PDEs to pricing American-type…

概率论 · 数学 2022-09-21 Elena Issoglio , Shuai Jing

We present a deep learning emulator for stochastic and chaotic spatio-temporal systems, explicitly conditioned on the parameter values of the underlying partial differential equations (PDEs). Our approach involves pre-training the model on…

机器学习 · 计算机科学 2025-09-12 Ira J. S. Shokar , Rich R. Kerswell , Peter H. Haynes

Inferring parameters of high-dimensional partial differential equations (PDEs) poses significant computational and inferential challenges, primarily due to the curse of dimensionality and the inherent limitations of traditional numerical…

计算工程、金融与科学 · 计算机科学 2025-09-18 Weihao Yan , Christoph Brune , Mengwu Guo

In the pursuit of autonomous spacecraft proximity maneuvers and docking(PMD), we introduce a novel Bayesian actor-critic reinforcement learning algorithm to learn a control policy with the stability guarantee. The PMD task is formulated as…

机器人学 · 计算机科学 2024-05-24 Desong Du , Naiming Qi , Yanfang Liu , Wei Pan

Deterministic policy gradient algorithms are foundational for actor-critic methods in controlling continuous systems, yet they often encounter inaccuracies due to their dependence on the derivative of the critic's value estimates with…

机器学习 · 计算机科学 2025-02-11 Baturay Saglam , Dionysis Kalogerias

Stochastic gradient descent (SGD), which updates the model parameters by adding a local gradient times a learning rate at each step, is widely used in model training of machine learning algorithms such as neural networks. It is observed…

机器学习 · 计算机科学 2017-06-01 Chang Xu , Tao Qin , Gang Wang , Tie-Yan Liu

Due to the curse of dimensionality, solving high dimensional parabolic partial differential equations (PDEs) has been a challenging problem for decades. Recently, a weak adversarial network (WAN) proposed in (Y.Zang et al., 2020) offered a…

数值分析 · 数学 2022-05-18 Paul Valsecchi Oliva , Yue Wu , Cuiyu He , Hao Ni

Despite the empirical success of the actor-critic algorithm, its theoretical understanding lags behind. In a broader context, actor-critic can be viewed as an online alternating update algorithm for bilevel optimization, whose convergence…

机器学习 · 计算机科学 2019-07-16 Zhuoran Yang , Yongxin Chen , Mingyi Hong , Zhaoran Wang

We consider a classical finite horizon optimal control problem for continuous-time pure jump Markov processes described by means of a rate transition measure depending on a control parameter and controlled by a feedback law. For this class…

概率论 · 数学 2015-01-20 Elena Bandini , Marco Fuhrman

Reinforcement learning (RL) and Deep Reinforcement Learning (DRL), in particular, have the potential to disrupt and are already changing the way we interact with the world. One of the key indicators of their applicability is their ability…

机器学习 · 计算机科学 2024-08-20 Nikolai Rozanov

We study a structured bi-level optimization problem where the upper-level objective is a smooth function and the lower-level problem is policy optimization in a Markov decision process (MDP). The upper-level decision variable parameterizes…

机器学习 · 计算机科学 2026-04-23 Sihan Zeng , Sujay Bhatt , Sumitra Ganesh , Alec Koppel

This paper proposes two efficient approximation methods to solve high-dimensional fully nonlinear partial differential equations (NPDEs) and second-order backward stochastic differential equations (2BSDEs), where such high-dimensional fully…

数值分析 · 数学 2023-01-18 Xu Xiao , Wenlin Qiu , Omid Nikan

In this paper, we propose a second-order deterministic actor-critic framework in reinforcement learning that extends the classical deterministic policy gradient method to exploit curvature information of the performance function. Building…

机器学习 · 计算机科学 2025-11-13 Arash Bahari Kordabad , Dean Brandner , Sebastien Gros , Sergio Lucia , Sadegh Soudjani

We construct the first rigorously justified probabilistic algorithm for recovering the solution operator of a hyperbolic partial differential equation (PDE) in two variables from input-output training pairs. The primary challenge of…

数值分析 · 数学 2026-02-03 Christopher Wang , Alex Townsend

In the present work we employ, for the first time, backward stochastic differential equations (BSDEs) to study the optimal control of semi-Markov processes on finite horizon, with general state and action spaces. More precisely, we prove…

最优化与控制 · 数学 2015-05-27 Elena Bandini , Fulvia Confortola

In this work, we generalize the reaction-diffusion equation in statistical physics, Schr\"odinger equation in quantum mechanics, Helmholtz equation in paraxial optics into the neural partial differential equations (NPDE), which can be…

机器学习 · 计算机科学 2024-10-11 Ping Guo , Kaizhu Huang , Zenglin Xu

Nonlinear stochastic differential equations (NSDEs) are a pillar of mathematical modeling for scientific and engineering applications. Accurate and efficient simulation of large-scale NSDEs is prohibitive on classical computers due to the…

量子物理 · 物理学 2026-03-16 Xiangyu Li , Ahmet Burak Catli , Ho Kiat Lim , Matthew Pocrnic , Dong An , Jin-Peng Liu , Nathan Wiebe

This paper is devoted to a stochastic differential game of functional forward-backward stochastic differential equation (FBSDE, for short). The associated upper and lower value functions of the stochastic differential game are defined by…

最优化与控制 · 数学 2013-01-03 Shaolin Ji , Qingmeng Wei

This paper addresses the problem of robust stabilization for linear hyperbolic Partial Differential Equations (PDEs) with Markov-jumping parameter uncertainty. We consider a 2 x 2 heterogeneous hyperbolic PDE and propose a control law using…

系统与控制 · 电气工程与系统科学 2026-03-13 Yihuai Zhang , Jean Auriol , Huan Yu