中文
相关论文

相关论文: Model-Free $\delta$-Policy Iteration Based on Damp…

200 篇论文

Consider the problem of approximating the optimal policy of a Markov decision process (MDP) by sampling state transitions. In contrast to existing reinforcement learning methods that are based on successive approximations to the nonlinear…

机器学习 · 计算机科学 2017-10-18 Mengdi Wang

This paper introduces Deep Policy Iteration (DPI), a novel approach that integrates the strengths of Neural Networks with the stability and convergence advantages of Policy Iteration (PI) to address high-dimensional stochastic Mean Field…

最优化与控制 · 数学 2024-07-15 Mouhcine Assouli , Badr Missaoui

This paper is concerned with the Proportional Integral (PI) regulation control of the left Neu-mann trace of a one-dimensional semilinear wave equation. The control input is selected as the right Neumann trace. The control design goes as…

最优化与控制 · 数学 2020-06-19 Hugo Lhachemi , Christophe Prieur , Emmanuel Trélat

Achieving real-time capability is an essential prerequisite for the industrial implementation of nonlinear model predictive control (NMPC). Data-driven model reduction offers a way to obtain low-order control models from complex digital…

系统与控制 · 电气工程与系统科学 2023-09-12 Jan C. Schulze , Danimir T. Doncevic , Nils Erwes , Alexander Mitsos

This paper presents an efficient model predictive path integral (MPPI) control framework for systems with complex nonlinear dynamics. To improve the computational efficiency of classic MPPI while preserving control performance, we replace…

机器人学 · 计算机科学 2026-03-06 Wenjian Hao , Yuxuan Fang , Zehui Lu , Shaoshuai Mou

The design of an automated vehicle controller can be generally formulated into an optimal control problem. This paper proposes a continuous-time finite-horizon approximate dynamicprogramming (ADP) method, which can synthesis off-line…

系统与控制 · 电气工程与系统科学 2020-07-07 Ziyu Lin , Jingliang Duan , Shengbo Eben Li , Haitong Ma , Yuming Yin

We provide a data-driven framework for optimal control of a continuous-time stochastic dynamical system. The proposed framework relies on the linear operator theory involving linear Perron-Frobenius (P-F) and Koopman operators. Our first…

最优化与控制 · 数学 2022-02-04 Umesh Vaidya , Duvan Tellez-Castro

Presented is a new method for calculating the time-optimal guidance control for a multiple vehicle pursuit-evasion system. A joint differential game of k pursuing vehicles relative to the evader is constructed, and a Hamilton-Jacobi-Isaacs…

最优化与控制 · 数学 2018-02-07 Matthew R. Kirchner , Robert Mar , Gary Hewer , Jérôme Darbon , Stanley Osher , Y. T. Chow

We formulate and analyze a new method for solving optimal control problems for systems governed by Volterra integral equations. Our method utilizes discretization of the original Volterra controlled system and a novel type of dynamic…

最优化与控制 · 数学 2007-05-23 S. A. Belbas

The paper studies a system of first order Hamilton-Jacobi equations with discontinuous coefficients, arising from a model of deterministic optimal debt management in infinite time horizon, with exponential discount and currency devaluation.…

最优化与控制 · 数学 2021-02-09 Antonio Marigonda , Khai T. Nguyen

This paper considers the distributed H-infinity leader-following tracking problem for a class of discrete time multi-agent systems with a high-dimensional dynamic leader. It is assumed that output information about the leader is only…

系统与控制 · 计算机科学 2013-09-03 Guanghui Wen , Valery Ugrinovskii

A data-based policy for iterative control task is presented. The proposed strategy is model-free and can be applied whenever safe input and state trajectories of a system performing an iterative task are available. These trajectories,…

系统与控制 · 计算机科学 2019-03-22 Ugo Rosolia , Xiaojing Zhang , Francesco Borrelli

We present a new methodology for studying non-Hamiltonian nonlinear systems based on an information theoretic extension of a renormalization group technique using a modified maximum entropy principle. We obtain a rigorous dimensionally…

计算物理 · 物理学 2013-06-28 M. Schmuck , M. Pradas , S. Kalliadasis , G. A. Pavliotis

This paper introduces the Hamilton-Jacobi-Bellman Proximal Policy Optimization (HJBPPO) algorithm into reinforcement learning. The Hamilton-Jacobi-Bellman (HJB) equation is used in control theory to evaluate the optimality of the value…

机器学习 · 计算机科学 2023-02-02 Amartya Mukherjee , Jun Liu

We propose a splitting approach to solve the second-order Hamilton--Jacobi equation, reducing it to a heat step and a purely first-order step. The latter is implemented using a gradient value policy iteration algorithm, enabling efficient…

最优化与控制 · 数学 2026-03-23 Alain Bensoussan , Thien P. B. Nguyen , Minh-Binh Tran , Son N. T. Tu

This article addresses the problem of data-driven numerical optimal control for unknown nonlinear systems. In our scenario, we suppose to have the possibility of performing multiple experiments (or simulations) on the system. Experiments…

系统与控制 · 电气工程与系统科学 2025-06-19 Marco Borghesi , Lorenzo Sforni , Giuseppe Notarstefano

Convex Q-learning is a recent approach to reinforcement learning, motivated by the possibility of a firmer theory for convergence, and the possibility of making use of greater a priori knowledge regarding policy or value function structure.…

最优化与控制 · 数学 2022-10-18 Fan Lu , Joel Mathias , Sean Meyn , Karanjit Kalsi

From the Hamilton-Jacobi-Bellman equation for the value function we derive a non-linear partial differential equation for the optimal portfolio strategy (the dynamic control). The equation is general in the sense that it does not depend on…

投资组合管理 · 定量金融 2013-11-20 Mads Nielsen

Constraint handling during tracking operations is at the core of many real-world control implementations and is well understood when dynamic models of the underlying system exist, yet becomes more challenging when data-driven models are…

系统与控制 · 电气工程与系统科学 2023-10-05 Ye Wang , Yujia Yang , Ye Pu , Chris Manzie

This paper studies an infinite horizon optimal tracking portfolio problem using capital injection in incomplete market models. The benchmark process is modelled by a geometric Brownian motion with zero drift driven by some unhedgeable risk.…

投资组合管理 · 定量金融 2024-11-01 Lijun Bo , Yijie Huang , Xiang Yu