中文
相关论文

相关论文: Policy Gradient for Continuous-Time Mean-Field Con…

200 篇论文

In this work, we consider policy-based methods for solving the reinforcement learning problem, and establish the sample complexity guarantees. A policy-based algorithm typically consists of an actor and a critic. We consider using various…

机器学习 · 计算机科学 2023-01-16 Zaiwei Chen , Siva Theja Maguluri

We study the global linear convergence of policy gradient (PG) methods for finite-horizon continuous-time exploratory linear-quadratic control (LQC) problems. The setting includes stochastic LQC problems with indefinite costs and allows…

最优化与控制 · 数学 2024-03-05 Michael Giegrich , Christoph Reisinger , Yufei Zhang

We present a unified framework for learning continuous control policies using backpropagation. It supports stochastic control by treating stochasticity in the Bellman equation as a deterministic function of exogenous noise. The product is a…

机器学习 · 计算机科学 2015-11-02 Nicolas Heess , Greg Wayne , David Silver , Timothy Lillicrap , Yuval Tassa , Tom Erez

We consider an optimal control problem where the average welfare of weakly interacting agents is of interest. We examine the mean-field control problem as the fluid approximation of the N-agent control problem with the setup of finite-state…

最优化与控制 · 数学 2024-02-13 Jingruo Sun

Reinforcement Learning (RL) has emerged as a powerful framework for sequential decision-making in dynamic environments, particularly when system parameters are unknown. This paper investigates RL-based control for entropy-regularized…

系统与控制 · 电气工程与系统科学 2025-12-02 Gabriel Diaz , Lucky Li , Wenhao Zhang

Model-free deep reinforcement learning has achieved great success in many domains, such as video games, recommendation systems and robotic control tasks. In continuous control tasks, widely used policies with Gaussian distributions results…

机器学习 · 计算机科学 2023-06-05 Lingwei Peng , Hui Qian , Zhebang Shen , Chao Zhang , Fei Li

We consider reinforcement learning (RL) methods for finding optimal policies in linear quadratic (LQ) mean field control (MFC) problems over an infinite horizon in continuous time, with common noise and entropy regularization. We study…

最优化与控制 · 数学 2024-08-06 Noufel Frikha , Huyên Pham , Xuanye Song

A dynamic mean field theory is developed for finite state and action Bayesian reinforcement learning in the large state space limit. In an analogy with statistical physics, the Bellman equation is studied as a disordered dynamical system;…

机器学习 · 统计学 2023-07-13 George Stamatescu

We study a high-dimensional stochastic optimization problem which features both control and stopping. In particular, a central planner steers a large population of particles, and can also remove particles at any time by paying a penalty. In…

最优化与控制 · 数学 2026-03-24 Pierre Cardaliaguet , Joe Jackson , Panagiotis E. Souganidis

We explore reinforcement learning methods for finding the optimal policy in the linear quadratic regulator (LQR) problem. In particular, we consider the convergence of policy gradient methods in the setting of known and unknown parameters.…

机器学习 · 计算机科学 2021-06-25 Ben Hambly , Renyuan Xu , Huining Yang

We propose a novel policy gradient method for multi-agent reinforcement learning, which leverages two different variance-reduction techniques and does not require large batches over iterations. Specifically, we propose a momentum-based…

Policy gradients in continuous control have been derived for both stochastic and deterministic policies. Here we study the relationship between the two. In a widely-used family of MDPs involving Gaussian control noise and quadratic control…

机器学习 · 计算机科学 2025-06-11 Emo Todorov

We introduce Group Policy Gradient (GPG), a family of critic-free policy-gradient estimators for general MDPs. Inspired by the success of GRPO's approach in Reinforcement Learning from Human Feedback (RLHF), GPG replaces a learned value…

机器学习 · 计算机科学 2025-10-07 Junhua Chen , Zixi Zhang , Hantao Zhong , Rika Antonova

We study interacting particle systems driven by noise, modeling phenomena such as opinion dynamics. We are interested in systems that exhibit phase transitions i.e. non-uniqueness of stationary states for the corresponding McKean-Vlasov…

最优化与控制 · 数学 2024-12-31 Sara Bicego , Dante Kalise , Grigorios A. Pavliotis

Deterministic policy gradient algorithms are foundational for actor-critic methods in controlling continuous systems, yet they often encounter inaccuracies due to their dependence on the derivative of the critic's value estimates with…

机器学习 · 计算机科学 2025-02-11 Baturay Saglam , Dionysis Kalogerias

Policy gradient methods are very attractive in reinforcement learning due to their model-free nature and convergence guarantees. These methods, however, suffer from high variance in gradient estimation, resulting in poor sample efficiency.…

机器学习 · 计算机科学 2018-11-16 Sergey Pankov

While the optimization landscape of policy gradient methods has been recently investigated for partially observed linear systems in terms of both static output feedback and dynamical controllers, they only provide convergence guarantees to…

最优化与控制 · 数学 2023-04-25 Feiran Zhao , Xingyun Fu , Keyou You

We study the optimal control of mean-field systems with heterogeneous and asymmetric interactions. This leads to considering a family of controlled Brownian diffusion processes with dynamics depending on the whole collection of marginal…

概率论 · 数学 2024-07-29 Anna De Crescenzo , Marco Fuhrman , Idris Kharroubi , Huyên Pham

The multidimensional Uncertain Volatility Model leads to robust option pricing problems under joint volatility and correlation uncertainty. Their numerical resolution quickly becomes challenging because the associated stochastic control…

We study mean-field control (MFC) problems with common noise using the control randomisation framework, where we substitute the control process with an independent Poisson point process, controlling its intensity instead. To address the…

最优化与控制 · 数学 2024-12-31 Robert Denkert , Idris Kharroubi , Huyên Pham