中文
相关论文

相关论文: Bridging Physics-Informed Neural Networks with Rei…

200 篇论文

We study reinforcement learning in hybrid discrete-continuous action spaces, such as settings where the discrete component selects a regime (or index) and the continuous component optimizes within it -- a structure common in robotics,…

机器学习 · 计算机科学 2026-05-15 Matias Alvo , Daniel Russo , Yash Kanoria

Proximal Policy Optimization (PPO) has become the predominant algorithm for on-policy reinforcement learning due to its scalability and empirical robustness across domains. However, there is a significant disconnect between the underlying…

Despite impressive results, reinforcement learning (RL) suffers from slow convergence and requires a large variety of tuning strategies. In this paper, we investigate the ability of RL algorithms on simple continuous control tasks. We show…

机器人学 · 计算机科学 2024-02-16 Daniel Layeghi , Steve Tonneau , Michael Mistry

Instability and slowness are two main problems in deep reinforcement learning. Even if proximal policy optimization (PPO) is the state of the art, it still suffers from these two problems. We introduce an improved algorithm based on…

机器学习 · 计算机科学 2019-10-01 Zhenyu Zhang , Xiangfeng Luo , Tong Liu , Shaorong Xie , Jianshu Wang , Wei Wang , Yang Li , Yan Peng

We propose a neural network approach for solving high-dimensional optimal control problems. In particular, we focus on multi-agent control problems with obstacle and collision avoidance. These problems immediately become high-dimensional,…

最优化与控制 · 数学 2022-05-05 Derek Onken , Levon Nurbekyan , Xingjian Li , Samy Wu Fung , Stanley Osher , Lars Ruthotto

In this paper, we propose a novel image restoration framework that integrates optimal control techniques with the Hamilton-Jacobi-Bellman (HJB) equation. Motivated by models from production planning, our method restores degraded images by…

偏微分方程分析 · 数学 2025-05-13 Dragos-Patru Covei

This paper studies the continuous-time reinforcement learning (RL) for optimal switching problems across multiple regimes. We consider a type of exploratory formulation under entropy regularization where the agent randomizes both the timing…

最优化与控制 · 数学 2025-12-23 Yijie Huang , Mengge Li , Xiang Yu , Zhou Zhou

Current reinforcement-learning methods are unable to directly learn policies that solve the minimum cost reach-avoid problem to minimize cumulative costs subject to the constraints of reaching the goal and avoiding unsafe states, as the…

机器学习 · 计算机科学 2024-10-31 Oswin So , Cheng Ge , Chuchu Fan

Reinforcement learning continuously optimizes decision-making based on real-time feedback reward signals through continuous interaction with the environment, demonstrating strong adaptive and self-learning capabilities. In recent years, it…

机器人学 · 计算机科学 2024-08-15 Zixiang Wang , Hao Yan , Yining Wang , Zhengjia Xu , Zhuoyue Wang , Zhizhong Wu

Recent literature has proposed approaches that learn control policies with high performance while maintaining safety guarantees. Synthesizing Hamilton-Jacobi (HJ) reachable sets has become an effective tool for verifying safety and…

系统与控制 · 电气工程与系统科学 2024-08-23 Milan Ganai , Sicun Gao , Sylvia Herbert

Proximal policy optimization (PPO) is one of the most successful deep reinforcement-learning methods, achieving state-of-the-art performance across a wide range of challenging tasks. However, its optimization behavior is still far from…

机器学习 · 计算机科学 2020-01-15 Yuhui Wang , Hao He , Chao Wen , Xiaoyang Tan

The purpose of this paper is to describe the numerical solution of the Hamilton-Jacobi-Bellman (HJB) for an optimal control problem for quantum spin systems. This HJB equation is a first order nonlinear partial differential equation defined…

量子物理 · 物理学 2011-10-05 Srinivas Sridharan , Matthew R. James

We propose a novel formulation for approximating reachable sets through a minimum discounted reward optimal control problem. The formulation yields a continuous solution that can be obtained by solving a Hamilton-Jacobi equation.…

最优化与控制 · 数学 2018-09-05 Anayo K. Akametalu , Shromona Ghosh , Jaime F. Fisac , Claire J. Tomlin

This paper first introduces a method to approximate the value function of high-dimensional optimal control by neural networks. Based on the established relationship between Pontryagin's maximum principle (PMP) and the value function of the…

最优化与控制 · 数学 2025-07-22 Mouhcine Assouli , Justina Gianatti , Badr Missaoui , Francisco J. Silva

This note lays part of the theoretical ground for a definition of differential systems modeling reinforcement learning in continuous time non-Markovian rough environments. Specifically we focus on optimal relaxed control of rough equations…

最优化与控制 · 数学 2024-02-29 Prakash Chakraborty , Harsha Honnappa , Samy Tindel

Proximal policy optimization (PPO) algorithm is a deep reinforcement learning algorithm with outstanding performance, especially in continuous control tasks. But the performance of this method is still affected by its exploration ability.…

机器学习 · 计算机科学 2020-11-12 Junwei Zhang , Zhenghao Zhang , Shuai Han , Shuai Lü

A gradient-enhanced functional tensor train cross approximation method for the resolution of the Hamilton-Jacobi-Bellman (HJB) equations associated to optimal feedback control of nonlinear dynamics is presented. The procedure uses samples…

数值分析 · 数学 2023-02-23 Sergey Dolgov , Dante Kalise , Luca Saluzzi

Solving the Hamilton-Jacobi-Bellman equation is important in many domains including control, robotics and economics. Especially for continuous control, solving this differential equation and its extension the Hamilton-Jacobi-Isaacs…

机器人学 · 计算机科学 2021-10-06 Michael Lutter , Boris Belousov , Shie Mannor , Dieter Fox , Animesh Garg , Jan Peters

We propose a variant of consensus-based optimization (CBO) algorithms, controlled-CBO, which introduces a feedback control term to improve convergence towards global minimizers of non-convex functions in multiple dimensions. The feedback…

最优化与控制 · 数学 2025-07-29 Yuyang Huang , Michael Herty , Dante Kalise , Nikolas Kantas

We study the problem of learning the optimal control policy for fine-tuning a given diffusion process, using general value function approximation. We develop a new class of algorithms by solving a variational inequality problem based on the…

机器学习 · 计算机科学 2025-09-03 Wenlong Mou