中文
相关论文

相关论文: Continuous-Time Fitted Value Iteration for Robust …

200 篇论文

An advantageous feature of piecewise constant policy timestepping for Hamilton-Jacobi-Bellman (HJB) equations is that different linear approximation schemes, and indeed different meshes, can be used for the resulting linear equations for…

数值分析 · 数学 2016-01-21 Christoph Reisinger , Peter Forsyth

We study the convergence rates of policy iteration (PI) for nonconvex viscous Hamilton--Jacobi equations using a discrete space-time scheme, where both space and time variables are discretized. We analyze the case with an uncontrolled…

数值分析 · 数学 2025-03-05 Xiaoqin Guo , Hung Vinh Tran , Yuming Paul Zhang

Stochastic optimal principle leads to the resolution of a partial differential equation (PDE), namely the Hamilton-Jacobi-Bellman (HJB) equation. In general, this equation cannot be solved analytically, thus numerical algorithms are the…

数值分析 · 数学 2021-09-14 Christelle Dleuna Nyoumbi , Antoine Tambue

Performative reinforcement learning is an emerging dynamical decision making framework, which extends reinforcement learning to the common applications where the agent's policy can change the environmental dynamics. Existing works on…

机器学习 · 计算机科学 2025-10-07 Ziyi Chen , Heng Huang

The Bellman equation and its continuous-time counterpart, the Hamilton-Jacobi-Bellman (HJB) equation, serve as necessary conditions for optimality in reinforcement learning and optimal control. While the value function is known to be the…

机器学习 · 计算机科学 2025-03-07 Haoxiang You , Lekan Molu , Ian Abraham

We provide a framework for incorporating robustness -- to perturbations in the transition dynamics which we refer to as model misspecification -- into continuous control Reinforcement Learning (RL) algorithms. We specifically focus on…

Controlling a non-statically bipedal robot is challenging due to the complex dynamics and multi-criterion optimization involved. Recent works have demonstrated the effectiveness of deep reinforcement learning (DRL) for simulation and…

机器人学 · 计算机科学 2021-12-23 Changxin Huang , Guangrun Wang , Zhibo Zhou , Ronghui Zhang , Liang Lin

Solving tasks in Reinforcement Learning is no easy feat. As the goal of the agent is to maximize the accumulated reward, it often learns to exploit loopholes and misspecifications in the reward signal resulting in unwanted behavior. While…

机器学习 · 计算机科学 2018-12-27 Chen Tessler , Daniel J. Mankowitz , Shie Mannor

Policy gradient methods are an appealing approach in reinforcement learning because they directly optimize the cumulative reward and can straightforwardly be used with nonlinear function approximators such as neural networks. The two main…

机器学习 · 计算机科学 2018-10-23 John Schulman , Philipp Moritz , Sergey Levine , Michael Jordan , Pieter Abbeel

This paper presents a physics-informed machine learning approach for synthesizing optimal feedback control policy for infinite-horizon optimal control problems by solving the Hamilton-Jacobi-Bellman (HJB) partial differential equation(PDE).…

系统与控制 · 电气工程与系统科学 2025-11-24 Tanay Raghunandan Srinivasa , Suraj Kumar

In recent years significant progress has been made in dealing with challenging problems using reinforcement learning.Despite its great success, reinforcement learning still faces challenge in continuous control tasks. Conventional methods…

机器学习 · 计算机科学 2020-02-04 Longxiang Shi , Shijian Li , Longbing Cao , Long Yang , Gang Zheng , Gang Pan

Reward shaping has been applied widely to accelerate Reinforcement Learning (RL) agents' training. However, a principled way of designing effective reward shaping functions, especially for complex continuous control problems, remains…

机器学习 · 计算机科学 2026-02-12 Mateo Juliani , Mingxuan Li , Elias Bareinboim

Robust control barrier functions (CBFs) provide a principled mechanism for smooth safety enforcement under worst-case disturbances. However, existing approaches typically rely on explicit, closed-form structure in the dynamics (e.g.,…

系统与控制 · 电气工程与系统科学 2026-04-16 Donggeon David Oh , Duy P. Nguyen , Haimin Hu , Jaime Fernández Fisac

We present a simple and easy to implement method for the numerical solution of a rather general class of Hamilton-Jacobi-Bellman (HJB) equations. In many cases, the considered problems have only a viscosity solution, to which, fortunately,…

计算金融 · 定量金融 2011-02-17 Jan Hendrik Witte , Christoph Reisinger

We study data-driven learning of robust stochastic control for infinite-horizon systems with potentially continuous state and action spaces. In many managerial settings--supply chains, finance, manufacturing, services, and dynamic…

机器学习 · 统计学 2025-11-18 Shengbo Wang , Jason Meng , Nian Si , Jose Blanchet , Zhengyuan Zhou

This paper presents a new fast and robust algorithm that provides fuel-optimal impulsive control input sequences that drive a linear time-variant system to a desired state at a specified time. This algorithm is applicable to a broad class…

最优化与控制 · 数学 2020-10-06 Adam W. Koenig , Simone D'Amico

We investigate the adaptive robust control framework for portfolio optimization and loss-based hedging under drift and volatility uncertainty. Adaptive robust problems offer many advantages but require handling a double optimization problem…

最优化与控制 · 数学 2020-05-06 Tao Chen , Michael Ludkovski

Reinforcement learning (RL) in continuous state-action spaces remains challenging in scientific computing due to poor sample efficiency and lack of pathwise physical consistency. We introduce Differential Reinforcement Learning…

机器学习 · 计算机科学 2026-02-06 Minh Nguyen , Chandrajit Bajaj

We consider a kind of stochastic exit time optimal control problems, in which the cost function is defined through a nonlinear backward stochastic differential equation. We study the regularity of the value function for such a control…

概率论 · 数学 2016-03-15 Rainer Buckdahn , Tianyang Nie

Unlike traditional model-based reinforcement learning approaches that estimate system parameters from data, non-model-based data-driven control learns the optimal policy directly from input-state data without any intermediate model…

最优化与控制 · 数学 2026-05-05 Leilei Cui , Zhong-Ping Jiang , Petter N. Kolm , Grégoire G. Macqueron
‹ 上一页 1 8 9 10 下一页 ›