English
Related papers

Related papers: A Novel Policy Iteration Algorithm for Nonlinear C…

200 papers

This paper introduces a novel model-free and a partially model-free algorithm for inverse optimal control (IOC), also known as inverse reinforcement learning (IRL), aimed at estimating the cost function of continuous-time nonlinear…

Systems and Control · Electrical Eng. & Systems 2025-03-20 Hamed Jabbari Asl , Eiji Uchibe

Considering that the decision-making environment faced by reinforcement learning (RL) agents is full of Knightian uncertainty, this paper describes the exploratory state dynamics equation in Knightian uncertainty to study the…

Optimization and Control · Mathematics 2026-01-27 Ziyu Li , Chen Fei , Weiyin Fei

Hamilton-Jacobi reachability (HJR) provides a value function that encodes the set of states from which a system with bounded control inputs can reach or avoid a target despite any bounded disturbance, and the corresponding robust, optimal…

Systems and Control · Electrical Eng. & Systems 2025-06-23 Will Sharpless , Yat Tin Chow , Sylvia Herbert

Stochastic optimal principle leads to the resolution of a partial differential equation (PDE), namely the Hamilton-Jacobi-Bellman (HJB) equation. In general, this equation cannot be solved analytically, thus numerical algorithms are the…

Numerical Analysis · Mathematics 2021-09-14 Christelle Dleuna Nyoumbi , Antoine Tambue

Learning from multi-step off-policy data collected by a set of policies is a core problem of reinforcement learning (RL). Approaches based on importance sampling (IS) often suffer from large variances due to products of IS ratios. Typical…

We propose a novel data-driven neural network (NN) optimization framework for solving an optimal stochastic control problem under stochastic constraints. Customized activation functions for the output layers of the NN are applied, which…

Optimization and Control · Mathematics 2023-06-21 Marc Chen , Mohammad Shirazi , Peter A. Forsyth , Yuying Li

We study a class of optimal control problems with state constraints where the state equation is a differential equation with delays. This class includes some problems arising in economics, in particular the so-called models with time to…

Optimization and Control · Mathematics 2009-07-09 Salvatore Federico , Ben Goldys , Fausto Gozzi

Reinforcement Learning (RL) offers a promising solution to enable evolutionary automated driving. However, the conventional RL method is always concerned with risk performance. The updated policy may not obtain a performance enhancement,…

Systems and Control · Electrical Eng. & Systems 2024-12-17 Jia Hu , Xuerun Yan , Tian Xu , Haoran Wang

This paper presents an inverse optimality method to solve the Hamilton-Jacobi-Bellman equation for a class of nonlinear problems for which the cost is quadratic and the dynamics are affine in the input. The method is inverse optimal because…

Optimization and Control · Mathematics 2011-10-11 Luis Rodrigues , Didier Henrion , Mehdi Abedinpour Fallah

We consider the infinite-horizon discounted optimal control problem formalized by Markov Decision Processes. We focus on several approximate variations of the Policy Iteration algorithm: Approximate Policy Iteration, Conservative Policy…

Artificial Intelligence · Computer Science 2014-05-13 Bruno Scherrer

Model predictive control (MPC) is widely used in process control due to its interpretability and ability to handle constraints. As a parametric policy in reinforcement learning (RL), MPC offers strong initial performance and low data…

Systems and Control · Electrical Eng. & Systems 2026-04-03 Dean Brandner , Sebastien Gros , Sergio Lucia

A control theoretic approach is presented in this paper for both batch and instantaneous updates of weights in feed-forward neural networks. The popular Hamilton-Jacobi-Bellman (HJB) equation has been used to generate an optimal weight…

Neural and Evolutionary Computing · Computer Science 2015-04-29 Vipul Arora , Laxmidhar Behera , Ajay Pratap Yadav

Designing a stabilizing controller for nonlinear systems is a challenging task, especially for high-dimensional problems with unknown dynamics. Traditional reinforcement learning algorithms applied to stabilization tasks tend to drive the…

Systems and Control · Electrical Eng. & Systems 2024-09-16 Thanin Quartz , Ruikun Zhou , Hans De Sterck , Jun Liu

This paper employs a policy iteration reinforcement learning (RL) method to study continuous-time linear-quadratic mean-field control problems in infinite horizon. The drift and diffusion terms in the dynamics involve the states, the…

Optimization and Control · Mathematics 2024-11-05 Na Li , Xun Li , Zuo Quan Xu

Policy gradient methods are an appealing approach in reinforcement learning because they directly optimize the cumulative reward and can straightforwardly be used with nonlinear function approximators such as neural networks. The two main…

Machine Learning · Computer Science 2018-10-23 John Schulman , Philipp Moritz , Sergey Levine , Michael Jordan , Pieter Abbeel

We propose a policy improvement algorithm for Reinforcement Learning (RL) which is called Rerouted Behavior Improvement (RBI). RBI is designed to take into account the evaluation errors of the Q-function. Such errors are common in RL when…

Machine Learning · Computer Science 2019-07-12 Elad Sarafian , Aviv Tamar , Sarit Kraus

We present a framework for efficient extraction of the viscosity solutions of nonlinear Hamilton-Jacobi equations with convex Hamiltonians. These viscosity solutions play a central role in areas such as front propagation, mean-field games,…

Quantum Physics · Physics 2026-02-17 Shi Jin , Nana Liu

We consider the problem of learning Nash equilibrial policies for two-player risk-sensitive collision-avoiding interactions. Solving the Hamilton-Jacobi-Isaacs equations of such general-sum differential games in real time is an open…

Robotics · Computer Science 2025-03-21 Lei Zhang , Siddharth Das , Tanner Merry , Wenlong Zhang , Yi Ren

We study the problem of learning the optimal control policy for fine-tuning a given diffusion process, using general value function approximation. We develop a new class of algorithms by solving a variational inequality problem based on the…

Machine Learning · Computer Science 2025-09-03 Wenlong Mou

In this paper, we demonstrate that policy iteration, introduced in the context of HJB equations in [Forsyth & Labahn, 2007], is an extremely simple generic algorithm for solving linear complementarity problems resulting from the finite…

Computational Finance · Quantitative Finance 2012-06-19 Christoph Reisinger , Jan Hendrik Witte