Related papers: Model-Free $\delta$-Policy Iteration Based on Damp…
Consider the problem of approximating the optimal policy of a Markov decision process (MDP) by sampling state transitions. In contrast to existing reinforcement learning methods that are based on successive approximations to the nonlinear…
This paper introduces Deep Policy Iteration (DPI), a novel approach that integrates the strengths of Neural Networks with the stability and convergence advantages of Policy Iteration (PI) to address high-dimensional stochastic Mean Field…
This paper is concerned with the Proportional Integral (PI) regulation control of the left Neu-mann trace of a one-dimensional semilinear wave equation. The control input is selected as the right Neumann trace. The control design goes as…
Achieving real-time capability is an essential prerequisite for the industrial implementation of nonlinear model predictive control (NMPC). Data-driven model reduction offers a way to obtain low-order control models from complex digital…
This paper presents an efficient model predictive path integral (MPPI) control framework for systems with complex nonlinear dynamics. To improve the computational efficiency of classic MPPI while preserving control performance, we replace…
The design of an automated vehicle controller can be generally formulated into an optimal control problem. This paper proposes a continuous-time finite-horizon approximate dynamicprogramming (ADP) method, which can synthesis off-line…
We provide a data-driven framework for optimal control of a continuous-time stochastic dynamical system. The proposed framework relies on the linear operator theory involving linear Perron-Frobenius (P-F) and Koopman operators. Our first…
Presented is a new method for calculating the time-optimal guidance control for a multiple vehicle pursuit-evasion system. A joint differential game of k pursuing vehicles relative to the evader is constructed, and a Hamilton-Jacobi-Isaacs…
We formulate and analyze a new method for solving optimal control problems for systems governed by Volterra integral equations. Our method utilizes discretization of the original Volterra controlled system and a novel type of dynamic…
The paper studies a system of first order Hamilton-Jacobi equations with discontinuous coefficients, arising from a model of deterministic optimal debt management in infinite time horizon, with exponential discount and currency devaluation.…
This paper considers the distributed H-infinity leader-following tracking problem for a class of discrete time multi-agent systems with a high-dimensional dynamic leader. It is assumed that output information about the leader is only…
A data-based policy for iterative control task is presented. The proposed strategy is model-free and can be applied whenever safe input and state trajectories of a system performing an iterative task are available. These trajectories,…
We present a new methodology for studying non-Hamiltonian nonlinear systems based on an information theoretic extension of a renormalization group technique using a modified maximum entropy principle. We obtain a rigorous dimensionally…
This paper introduces the Hamilton-Jacobi-Bellman Proximal Policy Optimization (HJBPPO) algorithm into reinforcement learning. The Hamilton-Jacobi-Bellman (HJB) equation is used in control theory to evaluate the optimality of the value…
We propose a splitting approach to solve the second-order Hamilton--Jacobi equation, reducing it to a heat step and a purely first-order step. The latter is implemented using a gradient value policy iteration algorithm, enabling efficient…
This article addresses the problem of data-driven numerical optimal control for unknown nonlinear systems. In our scenario, we suppose to have the possibility of performing multiple experiments (or simulations) on the system. Experiments…
Convex Q-learning is a recent approach to reinforcement learning, motivated by the possibility of a firmer theory for convergence, and the possibility of making use of greater a priori knowledge regarding policy or value function structure.…
From the Hamilton-Jacobi-Bellman equation for the value function we derive a non-linear partial differential equation for the optimal portfolio strategy (the dynamic control). The equation is general in the sense that it does not depend on…
Constraint handling during tracking operations is at the core of many real-world control implementations and is well understood when dynamic models of the underlying system exist, yet becomes more challenging when data-driven models are…
This paper studies an infinite horizon optimal tracking portfolio problem using capital injection in incomplete market models. The benchmark process is modelled by a geometric Brownian motion with zero drift driven by some unhedgeable risk.…