English
Related papers

Related papers: Beyond PID Controllers: PPO with Neuralized PID Po…

200 papers

Model inversion attacks pose a significant privacy risk by attempting to reconstruct private training data from trained models. Most of the existing methods either depend on gradient estimation or require white-box access to model…

Machine Learning · Computer Science 2025-02-21 Xinpeng Shou

On-policy reinforcement learning (RL), particularly Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO), has become the dominant paradigm for fine-tuning large language models (LLMs). While policy ratio clipping…

Machine Learning · Computer Science 2026-01-08 Yu Luo , Shuo Han , Yihan Hu , Dong Li , Jianye Hao

The Fermilab Proton Improvement Plan II, or PIP-II, would enable the world's most intense high-energy neutrino beam and would help scientists search for rare particle physics processes. The PIP-II goal is to deliver 1.2 MW of proton beam…

Accelerator Physics · Physics 2022-02-15 Sergei Nagaitsev , Valeri Lebedev

Continuous control of non-stationary environments is a major challenge for deep reinforcement learning algorithms. The time-dependency of the state transition dynamics aggravates the notorious stability problems of model-free deep…

Machine Learning · Computer Science 2025-11-05 Abdullah Akgül , Gulcin Baykal , Manuel Haußmann , Melih Kandemir

Reparameterization Policy Gradient (RPG) has emerged as a powerful paradigm for model-based reinforcement learning, enabling high sample efficiency by backpropagating gradients through differentiable dynamics. However, prior RPG approaches…

Machine Learning · Computer Science 2026-02-04 Hai Zhong , Zhuoran Li , Xun Wang , Longbo Huang

Drones are effective for reducing human activity and interactions by performing tasks such as exploring and inspecting new environments, monitoring resources and delivering packages. Drones need a controller to maintain stability and to…

Systems and Control · Electrical Eng. & Systems 2021-05-19 Azin Shamshirgaran , Hamed Javidi , Dan Simon

Assessing turbulence control effects for wall friction numerically is a significant challenge since it requires expensive simulations of turbulent fluid dynamics. We instead propose an efficient deep reinforcement learning (RL) framework…

Machine Learning · Computer Science 2025-10-07 Zelin Zhao , Zongyi Li , Kimia Hassibi , Kamyar Azizzadenesheli , Junchi Yan , H. Jane Bae , Di Zhou , Anima Anandkumar

Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO) are among the most successful policy gradient approaches in deep reinforcement learning (RL). While these methods achieve state-of-the-art performance across a…

Machine Learning · Computer Science 2020-06-22 Ahmed Touati , Amy Zhang , Joelle Pineau , Pascal Vincent

Multi-objective optimization (MOO) has been widely studied in literature because of its versatility in human-centered decision making in real-life applications. Recently, demand for dynamic MOO is fast-emerging due to tough market dynamics…

Artificial Intelligence · Computer Science 2026-04-14 Jiahuan Jin , Wenhao Zhao , Rong Qu , Jianfeng Ren , Xinan Chen , Qingfu Zhang , Ruibin Bai

Proximal Policy Optimization (PPO) is a ubiquitous on-policy reinforcement learning algorithm but is significantly less utilized than off-policy learning algorithms in multi-agent settings. This is often due to the belief that PPO is…

Machine Learning · Computer Science 2022-11-07 Chao Yu , Akash Velu , Eugene Vinitsky , Jiaxuan Gao , Yu Wang , Alexandre Bayen , Yi Wu

Fermilab is committed to upgrading its accelerator complex towards the intensity frontier to pursue HEP research in the neutrino sector and beyond. The upgrade has two steps: 1) the Proton Improvement Plan (PIP), which is underway, has its…

Accelerator Physics · Physics 2017-05-04 C. M. Bhat

For many applications of reinforcement learning it can be more convenient to specify both a reward function and constraints, rather than trying to design behavior through the reward function. For example, systems that physically interact…

Machine Learning · Computer Science 2017-05-31 Joshua Achiam , David Held , Aviv Tamar , Pieter Abbeel

We consider the problem of offline reinforcement learning with model-based control, whose goal is to learn a dynamics model from the experience replay and obtain a pessimism-oriented agent under the learned model. Current model-based…

Machine Learning · Computer Science 2021-09-16 Ruizhen Liu , Dazhi Zhong , Zhicong Chen

In this paper, we propose a novel framework for approximating the explicit MPC policy for linear parameter-varying systems using supervised learning. Our learning scheme guarantees feasibility and near-optimality of the approximated MPC…

Systems and Control · Electrical Eng. & Systems 2019-12-11 Xiaojing Zhang , Monimoy Bujarbaruah , Francesco Borrelli

Classical PID control is widely applied in an engineering system, with parameter regulation relying on a method like Trial - Error Tuning or the Ziegler - Nichols rule, mainly for a Single - Input Single - Output (SISO) system. However, the…

Systems and Control · Electrical Eng. & Systems 2025-04-22 Zimao Sheng , Hong'an Yang

Metallic coatings are used to enhance the durability of metal surfaces by protecting them from corrosion. These protective layers are typically deposited in a fluid state via a liquid film. Controlling instabilities in the liquid film is…

Fluid Dynamics · Physics 2026-04-29 Fabio Pino , Edoardo Fracchia , Benoit Scheid , Miguel A. Mendez

Sequential Bayesian optimal experimental design (SBOED) for PDE-governed inverse problems is computationally challenging, especially for infinite-dimensional random field parameters. High-fidelity approaches require repeated forward and…

Optimization and Control · Mathematics 2026-01-12 Kaichen Shen , Peng Chen

In this paper, we propose a new algorithm PPG (Proximal Policy Gradient), which is close to both VPG (vanilla policy gradient) and PPO (proximal policy optimization). The PPG objective is a partial variation of the VPG objective and the…

Machine Learning · Computer Science 2020-10-21 Ju-Seung Byun , Byungmoon Kim , Huamin Wang

PID control architectures are widely used in industrial applications. Despite their low number of open parameters, tuning multiple, coupled PID controllers can become tedious in practice. In this paper, we extend PILCO, a model-based policy…

Machine Learning · Computer Science 2017-03-09 Andreas Doerr , Duy Nguyen-Tuong , Alonso Marco , Stefan Schaal , Sebastian Trimpe

Many real-world problems require sequential decisions under uncertainty: when to inject or withdraw gas from storage, how to rebalance a pension portfolio each month, what temperature profile to run through a pharmaceutical reactor chain.…

Machine Learning · Computer Science 2026-05-08 Dmitri Goloubentsev , Natalija Karpichina