中文
相关论文

相关论文: Probabilistic Framework of Howard's Policy Iterati…

200 篇论文

Parameter inference for stochastic differential equation mixed effects models (SDEMEMs) is a challenging problem. Analytical solutions for these models are rarely available, which means that the likelihood is also intractable. In this case,…

统计计算 · 统计学 2019-09-30 Imke Botha , Robert Kohn , Christopher Drovandi

We propose the Compound BSDE method, a fully forward, deep-learning-based approach for solving a broad class of problems in financial mathematics, including optimal stopping. The method is based on a reformulation of option pricing problems…

计算金融 · 定量金融 2026-02-02 Zhipeng Huang , Cornelis W. Oosterlee

The paper is directly motivated by the pricing of vulnerable European and American options in a general hazard process setup and a related study of the corresponding pre-default backward stochastic differential equations (BSDE) and…

概率论 · 数学 2022-12-27 Libo Li , Ruyi Liu , Marek Rutkowski

We propose some numerical schemes for forward-backward stochastic differential equations (FBSDEs) based on a new fundamental concept of transposition solutions. These schemes exploit time-splitting methods for the variation of constants…

数值分析 · 数学 2018-05-01 Kazufumi Ito , Yufei Zhang , Jun Zou

Probabilistic Manifold Decomposition (PMD)\cite{doi:10.1137/25M1738863}, developed in our earlier work, provides a nonlinear model reduction by embedding high-dimensional dynamics onto low-dimensional probabilistic manifolds. The PMD has…

数值分析 · 数学 2026-01-13 Jiaming Guo , Dunhui Xiao

The so-called block-term decomposition (BTD) tensor model, especially in its rank-$(L_r,L_r,1)$ version, has been recently receiving increasing attention due to its enhanced ability of representing systems and signals that are composed of…

统计方法学 · 统计学 2022-05-04 Paris V. Giampouras , Athanasios A. Rontogiannis , Eleftherios Kofidis

We consider inexact policy iteration methods for large-scale infinite-horizon discounted MDPs with finite spaces, a variant of policy iteration where the policy evaluation step is implemented inexactly using an iterative solver for linear…

最优化与控制 · 数学 2024-04-10 Matilde Gargiani , Robin Sieber , Efe Balta , Dominic Liao-McPherson , John Lygeros

The Bradley-Terry model is a popular approach to describe probabilities of the possible outcomes when elements of a set are repeatedly compared with one another in pairs. It has found many applications including animal behaviour, chess…

统计方法学 · 统计学 2015-03-17 Francois Caron , Arnaud Doucet

In the realm of statistical learning, the increasing volume of accessible data and increasing model complexity necessitate robust methodologies. This paper explores two branches of robust Bayesian methods in response to this trend. The…

统计方法学 · 统计学 2024-12-02 Masahiro Tanaka

There has been substantial progress in the inference of formal behavioural specifications from sample trajectories, for example, using Linear Temporal Logic (LTL). However, these techniques cannot handle specifications that correctly…

计算机科学中的逻辑 · 计算机科学 2025-05-20 Rajarshi Roy , Yash Pote , David Parker , Marta Kwiatkowska

Model-based reinforcement learning algorithms tend to achieve higher sample efficiency than model-free methods. However, due to the inevitable errors of learned models, model-based methods struggle to achieve the same asymptotic performance…

机器学习 · 计算机科学 2019-12-02 Qi Zhou , Houqiang Li , Jie Wang

We consider off-policy temporal-difference (TD) learning methods for policy evaluation in Markov decision processes with finite spaces and discounted reward criteria, and we present a collection of convergence results for several…

机器学习 · 计算机科学 2018-03-30 Huizhen Yu

This article introduces a framework for measuring the uncertain behaviour of a changing system in terms of the solution of a class of fractional stochastic differential equations (fsDEs). This is accomplished via operational matrices based…

综合数学 · 数学 2025-06-03 O. T. Birgani , J. F. Peters , S. Kouhkani

The convergence of deterministic policy gradient under the Hadamard parameterization is studied in the tabular setting and the linear convergence of the algorithm is established. To this end, we first show that the error decreases at an…

最优化与控制 · 数学 2023-11-28 Jiacai Liu , Jinchi Chen , Ke Wei

Optimism about the poorly understood states and actions is the main driving force of exploration for many provably-efficient reinforcement learning algorithms. We propose optimism in the face of sensible value functions (OFVF)- a novel…

机器学习 · 计算机科学 2019-04-19 Reazul H. Russel , Tianyi Gu , Marek Petrik

This paper considers a non-Markov control problem arising in a financial market where asset returns depend on hidden factors. The problem is non-Markov because nonlinear filtering is required to make inference on these factors, and hence…

数理金融 · 定量金融 2018-07-24 Andrew Papanicolaou

Forward-backward selection is one of the most basic and commonly-used feature selection algorithms available. It is also general and conceptually applicable to many different types of data. In this paper, we propose a heuristic that…

机器学习 · 计算机科学 2017-05-31 Giorgos Borboudakis , Ioannis Tsamardinos

The semi-parametric Cox proportional hazards regression model has been widely used for many years in several applied sciences. However, a fully parametric proportional hazards model, if appropriately assumed, can often lead to more…

统计方法学 · 统计学 2020-09-29 Amarnath Nandy , Abhik Ghosh , Ayanendranath Basu , Leandro Pardo

Novel multi-step predictor-corrector numerical schemes have been derived for approximating decoupled forward-backward stochastic differential equations (FBSDEs). The stability and high order rate of convergence of the schemes are rigorously…

数值分析 · 数学 2021-02-12 Qiang Han , Shaolin Ji

Mean-field backward doubly stochastic differential equations (MF-BDSDEs, for short) are introduced and studied. The existence and uniqueness of solutions for MF-BDSDEs is established. One probabilistic interpretation for the solutions to a…

概率论 · 数学 2011-08-30 Tianxiao Wang , Qingfeng Zhu , Yufeng Shi