中文
相关论文

相关论文: Continuous-Time Fitted Value Iteration for Robust …

200 篇论文

We study the policy iteration algorithm (PIA) for entropy-regularized stochastic control problems on an infinite time horizon with a large discount rate, focusing on two main scenarios. First, we analyze PIA with bounded coefficients where…

最优化与控制 · 数学 2025-05-28 Hung Vinh Tran , Zhenhua Wang , Yuming Paul Zhang

Maximum entropy reinforcement learning (RL) methods have been successfully applied to a range of challenging sequential decision-making and control tasks. However, most of existing techniques are designed for discrete-time systems. As a…

最优化与控制 · 数学 2020-09-29 Jeongho Kim , Insoon Yang

Safety is the priority concern when applying reinforcement learning (RL) algorithms to real-world control problems. While policy iteration provides a fundamental algorithm for standard RL, an analogous theoretical algorithm for safe RL…

机器学习 · 计算机科学 2025-03-14 Yujie Yang , Zhilong Zheng , Shengbo Eben Li , Wei Xu , Jingjing Liu , Xianyuan Zhan , Ya-Qin Zhang

In this paper, we introduce Hamilton-Jacobi-Bellman (HJB) equations for Q-functions in continuous time optimal control problems with Lipschitz continuous controls. The standard Q-function used in reinforcement learning is shown to be the…

最优化与控制 · 数学 2020-05-05 Jeongho Kim , Insoon Yang

We consider continuous-time stochastic optimal control problems featuring Conditional Value-at-Risk (CVaR) in the objective. The major difficulty in these problems arises from time-inconsistency, which prevents us from directly using…

最优化与控制 · 数学 2020-05-27 Christopher W. Miller , Insoon Yang

In this paper, we propose a new policy iteration algorithm to compute the value function and the optimal controls of continuous time stochastic control problems. The algorithm relies on successive approximations using linear-quadratic…

最优化与控制 · 数学 2024-09-09 Dylan Possamaï , Ludovic Tangpi

The uncertainties in plant dynamics remain a challenge for nonlinear control problems. This paper develops a ternary policy iteration (TPI) algorithm for solving nonlinear robust control problems with bounded uncertainties. The controller…

系统与控制 · 电气工程与系统科学 2020-07-15 Jie Li , Shengbo Eben Li , Yang Guan , Jingliang Duan , Wenyu Li , Yuming Yin

We propose a novel formulation for approximating reachable sets through a minimum discounted reward optimal control problem. The formulation yields a continuous solution that can be obtained by solving a Hamilton-Jacobi equation.…

最优化与控制 · 数学 2018-09-05 Anayo K. Akametalu , Shromona Ghosh , Jaime F. Fisac , Claire J. Tomlin

This paper investigates the optimal control problems for the finite-horizon continuous-time Markov decision processes with delay-dependent control policies. We develop compactification methods in decision processes, and show that the…

概率论 · 数学 2023-07-06 Zhong-Wei Liao , Jinghai Shao

Recent focus on robustness to adversarial attacks for deep neural networks produced a large variety of algorithms for training robust models. Most of the effective algorithms involve solving the min-max optimization problem for training…

机器学习 · 计算机科学 2021-03-03 Yasaman Esfandiari , Aditya Balu , Keivan Ebrahimi , Umesh Vaidya , Nicola Elia , Soumik Sarkar

This paper is concerned with an optimal control problem for a forward-backward stochastic differential equation (FBSDE, for short) with a recursive cost functional determined by a backward stochastic Volterra integral equation (BSVIE, for…

最优化与控制 · 数学 2022-09-20 Hanxiao Wang , Jiongmin Yong , Chao Zhou

We study the problem of learning policies that maximize cumulative reward while satisfying safety constraints, even when the real environment differs from a simulator or nominal model. We focus on robust constrained Markov decision…

机器学习 · 计算机科学 2025-11-12 Sourav Ganguly , Arnob Ghosh

An optimal control problem is considered for a stochastic differential equation with the cost functional determined by a backward stochastic Volterra integral equation (BSVIE, for short). This kind of cost functional can cover the general…

最优化与控制 · 数学 2019-11-13 Hanxiao Wang , Jiongmin Yong

This paper addresses the problem of utility maximization under uncertain parameters. In contrast with the classical approach, where the parameters of the model evolve freely within a given range, we constrain them via a penalty function. We…

最优化与控制 · 数学 2022-03-08 Ivan Guo , Nicolas Langrené , Grégoire Loeper , Wei Ning

We treat infinite horizon optimal control problems by solving the associated stationary Hamilton-Jacobi-Bellman (HJB) equation numerically to compute the value function and an optimal feedback law. The dynamical systems under consideration…

最优化与控制 · 数学 2021-05-19 Mathias Oster , Leon Sallandt , Reinhold Schneider

Robust Reinforcement Learning tries to make predictions more robust to changes in the dynamics or rewards of the system. This problem is particularly important when the dynamics and rewards of the environment are estimated from the data. In…

机器学习 · 计算机科学 2022-06-15 Pierre Clavier , Stéphanie Allassonière , Erwan Le Pennec

We propose a mesh-free policy iteration framework that combines classical dynamic programming with physics-informed neural networks (PINNs) to solve high-dimensional, nonconvex Hamilton--Jacobi--Isaacs (HJI) equations arising in stochastic…

数值分析 · 数学 2025-07-24 Hee Jun Yang , Minjung Gim , Yeoneung Kim

We consider a Bolza-type optimal control problem for a dynamical system described by a fractional differential equation with the Caputo derivative of an order $\alpha \in (0, 1)$. The value of this problem is introduced as a functional in a…

最优化与控制 · 数学 2019-08-06 Mikhail I. Gomoyunov

The goal of robust reinforcement learning (RL) is to learn a policy that is robust against the uncertainty in model parameters. Parameter uncertainty commonly occurs in many real-world RL applications due to simulator modeling errors,…

机器学习 · 计算机科学 2022-10-19 Kishan Panaganti , Zaiyan Xu , Dileep Kalathil , Mohammad Ghavamzadeh

In this paper, we introduce a model-based deep-learning approach to solve finite-horizon continuous-time stochastic control problems with jumps. We iteratively train two neural networks: one to represent the optimal policy and the other to…

机器学习 · 计算机科学 2026-01-16 Patrick Cheridito , Jean-Loup Dupret , Donatien Hainaut