中文
相关论文

相关论文: Probabilistic Framework of Howard's Policy Iterati…

200 篇论文

In this paper, we introduce a new kind of reflected backward stochastic differential equations (RBSDEs) driven by a martingale, in a Markov chain model, but not driven by Brownian motion, and give existence and uniqueness results for the…

概率论 · 数学 2015-05-14 Dimbinirina Ramarimbahoaka , Zhe Yang , Robert J. Elliott

We propose a deep learning algorithm for solving high-dimensional parabolic integro-differential equations (PIDEs) and high-dimensional forward-backward stochastic differential equations with jumps (FBSDEJs), where the jump-diffusion…

数值分析 · 数学 2023-01-31 Wansheng Wang , Jie Wang , Jinping Li , Feifei Gao , Yi Fu

Policy iteration is one of the classical frameworks of reinforcement learning, which requires a known initial stabilizing control. However, finding the initial stabilizing control depends on the known system model. To relax this requirement…

系统与控制 · 电气工程与系统科学 2025-03-20 Dongdong Li , Jiuxiang Dong

Recently, the deep learning method has been used for solving forward-backward stochastic differential equations (FBSDEs) and parabolic partial differential equations (PDEs). It has good accuracy and performance for high-dimensional…

数值分析 · 数学 2020-02-04 Shaolin Ji , Shige Peng , Ying Peng , Xichuan Zhang

The computation of Bayesian estimates of system parameters and functions of them on the basis of observed system performance data is a common problem within system identification. This is a previously studied issue where stochastic…

统计计算 · 统计学 2018-05-09 Johan Dahlin , Adrian Wills , Brett Ninness

Backward stochastic differential equations (BSDEs) appear in numeruous applications. Classical approximation methods suffer from the curse of dimensionality and deep learning-based approximation methods are not known to converge to the BSDE…

概率论 · 数学 2022-04-20 Martin Hutzenthaler , Tuan Anh Nguyen

We propose a numerical method for solving high dimensional fully nonlinear partial differential equations (PDEs). Our algorithm estimates simultaneously by backward time induction the solution and its gradient by multi-layer neural…

最优化与控制 · 数学 2021-01-27 Huyen Pham , Xavier Warin , Maximilien Germain

This paper introduces a loss-based generalized Bayesian methodology for high-dimensional robust regression with serially correlated errors and predictors. The proposed framework employs a novel scaled pseudo-Huber (SPH) loss function, which…

统计方法学 · 统计学 2025-03-13 Saptarshi Chakraborty , Kshitij Khare , George Michailidis

We present a novel numerical method for solving McKean-Vlasov forward-backward stochastic differential equations (MV-FBSDEs) with common noise, combining Picard iterations, elicitability and deep learning. The key innovation involves…

机器学习 · 计算机科学 2025-12-18 Felipe J. P. Antunes , Yuri F. Saporito , Sebastian Jaimungal

We propose a deep signature/log-signature FBSDE algorithm to solve forward-backward stochastic differential equations (FBSDEs) with state and path dependent features. By incorporating the deep signature/log-signature transformation into the…

机器学习 · 计算机科学 2022-08-22 Qi Feng , Man Luo , Zhaoyu Zhang

Partially-observed Boolean dynamical systems (POBDS) are a general class of nonlinear models with application in estimation and control of Boolean processes based on noisy and incomplete measurements. The optimal minimum mean square error…

统计方法学 · 统计学 2017-03-08 Mahdi Imani , Ulisses Braga-Neto

Block coordinate descent (BCD) methods are prevalent in large scale optimization problems due to the low memory and computational costs per iteration, the predisposition to parallelization, and the ability to exploit the structure of the…

最优化与控制 · 数学 2025-10-31 Luis Briceño-Arias , Paulo Gonçalves , Guillaume Lauga , Nelly Pustelnik , Elisa Riccietti

In the theory of Partially Observed Markov Decision Processes (POMDPs), existence of optimal policies have in general been established via converting the original partially observed stochastic control problem to a fully observed one on the…

最优化与控制 · 数学 2022-01-11 Ali Devran Kara , Serdar Yuksel

We propose a new policy gradient method, named homotopic policy mirror descent (HPMD), for solving discounted, infinite horizon MDPs with finite state and action spaces. HPMD performs a mirror descent type policy update with an additional…

机器学习 · 计算机科学 2022-11-30 Yan Li , Guanghui Lan , Tuo Zhao

Existing score-based methods for inverse problems often resort to approximate minimization of the KL divergence between the inversion distribution and the Bayesian posterior. Such an approximation leads to severe mode collapse and…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Weimin Bai , Yuxuan Gu , Yifei Wang , Weijian Luo , He Sun

Outliers can seriously distort statistical inference by inducing excessive sensitivity in the likelihood function, thereby compromising the reliability of Bayesian estimation. To address this issue, we develop a robust Bayesian estimation…

统计理论 · 数学 2026-02-09 Jeongho Lee , Junmo Song

Regularized MDPs serve as a smooth version of original MDPs. However, biased optimal policy always exists for regularized MDPs. Instead of making the coefficient{\lambda}of regularized term sufficiently small, we propose an adaptive…

机器学习 · 计算机科学 2020-11-03 Wenhao Yang , Xiang Li , Guangzeng Xie , Zhihua Zhang

In this paper, we aim to solve the high dimensional stochastic optimal control problem from the view of the stochastic maximum principle via deep learning. By introducing the extended Hamiltonian system which is essentially an FBSDE with a…

最优化与控制 · 数学 2021-06-23 Shaolin Ji , Shige Peng , Ying Peng , Xichuan Zhang

This paper analyses the forecasting performance of a new class of factor models with martingale difference errors (FMMDE) recently introduced by Lee and Shao (2018). The FMMDE makes it possible to retrieve a transformation of the original…

计量经济学 · 经济学 2023-06-23 Luca Mattia Rolla , Alessandro Giovannelli

We study the convergence rate of stochastic optimization of exact (NP-hard) objectives, for which only biased estimates of the gradient are available. We motivate this problem in the context of learning the structure and parameters of Ising…

机器学习 · 计算机科学 2018-11-16 Jean Honorio