中文
相关论文

相关论文: Probabilistic Framework of Howard's Policy Iterati…

200 篇论文

Autonomous systems are often required to operate in partially observable environments. They must reliably execute a specified objective even with incomplete information about the state of the environment. We propose a methodology to…

人工智能 · 计算机科学 2020-01-14 Maxime Bouton , Jana Tumova , Mykel J. Kochenderfer

The present paper studies a kind of robust optimization problems with constraint. The problem is formulated through Backward Stochastic Differential Equations (BSDEs) with quadratic generators. A necessary condition is established for the…

最优化与控制 · 数学 2024-02-14 Peng Luo , Alexander Schied , Xiaole Xue

The majority of machine learning methods can be regarded as the minimization of an unavailable risk function. To optimize the latter, given samples provided in a streaming fashion, we define a general stochastic Newton algorithm and its…

统计理论 · 数学 2023-06-30 Claire Boyer , Antoine Godichon-Baggioni

We introduce Bayesian least-squares policy iteration (BLSPI), an off-policy, model-free, policy iteration algorithm that uses the Bayesian least-squares temporal-difference (BLSTD) learning algorithm to evaluate policies. An online variant…

机器学习 · 计算机科学 2019-04-09 Nikolaos Tziortziotis , Christos Dimitrakakis , Michalis Vazirgiannis

Policy evaluation is a crucial step in many reinforcement-learning procedures, which estimates a value function that predicts states' long-term value under a given policy. In this paper, we focus on policy evaluation with linear function…

机器学习 · 计算机科学 2017-06-12 Simon S. Du , Jianshu Chen , Lihong Li , Lin Xiao , Dengyong Zhou

In this paper we discuss policy iteration methods for approximate solution of a finite-state discounted Markov decision problem, with a focus on feature-based aggregation methods and their connection with deep reinforcement learning…

机器学习 · 计算机科学 2018-08-23 Dimitri P. Bertsekas

We investigate the use of the Multiple Optimised Parameter Estimation and Data compression algorithm (MOPED) for data compression and faster evaluation of likelihood functions. Since MOPED only guarantees maintaining the Fisher matrix of…

天体物理仪器与方法 · 物理学 2011-05-17 Philip Graff , Mike Hobson , Anthony Lasenby

The partitioning of data for estimation and calibration critically impacts the performance of propensity score based estimators like inverse probability weighting (IPW) and double/debiased machine learning (DML) frameworks. We extend recent…

机器学习 · 统计学 2025-05-20 Sven Klaassen , Jan Rabenseifner , Jannis Kueck , Philipp Bach

This work develops novel strategies for optimal planning with semantic observations using continuous state partially observable markov decision processes (CPOMDPs). Two major innovations are presented in relation to Gaussian mixture (GM)…

人工智能 · 计算机科学 2019-08-09 Luke Burks , Ian Loefgren , Nisar Ahmed

Addressing uncertainty is critical for autonomous systems to robustly adapt to the real world. We formulate the problem of model uncertainty as a continuous Bayes-Adaptive Markov Decision Process (BAMDP), where an agent maintains a…

机器人学 · 计算机科学 2019-05-09 Gilwoo Lee , Brian Hou , Aditya Mandalika , Jeongseok Lee , Sanjiban Choudhury , Siddhartha S. Srinivasa

In [5] the authors obtained Mean-Field backward stochastic differential equations (BSDE) associated with a Mean-field stochastic differential equation (SDE) in a natural way as limit of some highly dimensional system of forward and backward…

概率论 · 数学 2007-11-21 Rainer Buckdahn , Juan Li , Shige Peng

In this work, we study the numerical approximation of a class of singular fully coupled forward backward stochastic differential equations. These equations have a degenerate forward component and non-smooth terminal condition. They are…

数值分析 · 数学 2022-08-17 Jean-François Chassagneux , Mohan Yang

In this paper, we study the connections between three concepts - the reverse H\"older inequality for matrix-valued martingales, the well-posedness of linear BSDEs with unbounded coefficients, and the well-posedness of quadratic BSDE…

概率论 · 数学 2022-03-01 Joe Jackson

Constrained decision-making is essential for designing safe policies in real-world control systems, yet simulated environments often fail to capture real-world adversities. We consider the problem of learning a policy that will maximize the…

机器学习 · 计算机科学 2026-02-10 Sourav Ganguly , Kishan Panaganti , Arnob Ghosh , Adam Wierman

The famous Policy Iteration algorithm alternates between policy improvement and policy evaluation. Implementations of this algorithm with several variants of the latter evaluation stage, e.g, $n$-step and trace-based returns, have been…

人工智能 · 计算机科学 2018-08-01 Yonathan Efroni , Gal Dalal , Bruno Scherrer , Shie Mannor

This is one of our series papers on multistep schemes for solving forward backward stochastic differential equations (FBSDEs) and related problems. Here we extend (with non-trivial updates) our multistep schemes in [W. Zhao, Y. Fu and T.…

数值分析 · 数学 2015-02-12 Kong Tao , Weidong Zhao , Tao Zhou

Forward-backward stochastic differential equations (FBSDEs) have attracted significant attention since they were introduced almost 30 years ago, due to their wide range of applications, from solving non-linear PDEs to pricing American-type…

概率论 · 数学 2022-09-21 Elena Issoglio , Shuai Jing

We propose a novel framework for solving a class of Partial Integro-Differential Equations (PIDEs) and Forward-Backward Stochastic Differential Equations with Jumps (FBSDEJs) through a deep learning-based approach. This method, termed the…

数值分析 · 数学 2024-12-17 Zaijun Ye , Wansheng Wang

Gradient-based approaches to direct policy search in reinforcement learning have received much recent attention as a means to solve problems of partial observability and to avoid some of the problems associated with policy degradation in…

人工智能 · 计算机科学 2019-11-18 Jonathan Baxter , Peter L. Bartlett

Bayesian model comparison (BMC) offers a principled approach for assessing the relative merits of competing computational models and propagating uncertainty into model selection decisions. However, BMC is often intractable for the popular…

机器学习 · 统计学 2023-11-27 Lasse Elsemüller , Martin Schnuerch , Paul-Christian Bürkner , Stefan T. Radev