中文
相关论文

相关论文: Policy Optimization for Continuous-time Linear-Qua…

200 篇论文

Convergence of the policy iteration method for discrete and continuous optimal control problems holds under general assumptions. Moreover, in some circumstances, it is also possible to show a quadratic rate of convergence for the algorithm.…

最优化与控制 · 数学 2022-03-02 Fabio Camilli , Qing Tang

In this paper, we study the global convergence of model-based and model-free policy gradient descent and natural policy gradient descent algorithms for linear quadratic deep structured teams. In such systems, agents are partitioned into a…

多智能体系统 · 计算机科学 2020-12-16 Vida Fathi , Jalal Arabneydi , Amir G. Aghdam

In this paper, we present a unifying framework for analyzing equilibria and designing interventions for large network games sampled from a stochastic network formation process represented by a graphon. We first introduce a new class of…

计算机科学与博弈论 · 计算机科学 2020-07-01 Francesca Parise , Asuman Ozdaglar

Linear quadratic graphon field games (LQ-GFGs) are defined to be LQ games which involve a large number of agents that are weakly coupled via a weighted undirected graph on which each node represents an agent. The links of the graph…

系统与控制 · 电气工程与系统科学 2021-06-24 Shuang Gao , Rinel Foguen Tchuendom , Peter E. Caines

In this paper, zero-sum mean-field type games (ZSMFTG) with linear dynamics and quadratic cost are studied under infinite-horizon discounted utility function. ZSMFTG are a class of games in which two decision makers whose utilities sum to…

最优化与控制 · 数学 2020-09-02 René Carmona , Kenza Hamidouche , Mathieu Laurière , Zongjun Tan

The emergence of the graphon theory of large networks and their infinite limits has enabled the formulation of a theory of the centralized control of dynamical systems distributed on asymptotically infinite networks (Gao and Caines, IEEE…

最优化与控制 · 数学 2021-12-30 Peter E. Caines , Minyi Huang

Policy optimization methods with function approximation are widely used in multi-agent reinforcement learning. However, it remains elusive how to design such algorithms with statistical guarantees. Leveraging a multi-agent performance…

机器学习 · 计算机科学 2023-05-09 Yulai Zhao , Zhuoran Yang , Zhaoran Wang , Jason D. Lee

Entropy regularization has been extensively adopted to improve the efficiency, the stability, and the convergence of algorithms in reinforcement learning. This paper analyzes both quantitatively and qualitatively the impact of entropy…

最优化与控制 · 数学 2021-12-10 Xin Guo , Renyuan Xu , Thaleia Zariphopoulou

We present a Reinforcement Learning (RL) algorithm to solve infinite horizon asymptotic Mean Field Game (MFG) and Mean Field Control (MFC) problems. Our approach can be described as a unified two-timescale Mean Field Q-learning: The…

最优化与控制 · 数学 2021-06-01 Andrea Angiuli , Jean-Pierre Fouque , Mathieu Laurière

Optimization of parameterized policies for reinforcement learning (RL) is an important and challenging problem in artificial intelligence. Among the most common approaches are algorithms based on gradient ascent of a score function…

This paper analyzes a class of infinite-time-horizon stochastic games with singular controls motivated from the partially reversible problem. It provides an explicit solution for the mean-field game (MFG) and presents sensitivity analysis…

最优化与控制 · 数学 2020-08-12 Haoyang Cao , Xin Guo

Multi-agent robust reinforcement learning, also known as multi-player robust Markov games (RMGs), is a crucial framework for modeling competitive interactions under environmental uncertainties, with wide applications in multi-agent systems.…

机器学习 · 计算机科学 2024-12-31 Yuchen Jiao , Gen Li

Modern robotic systems frequently engage in complex multi-agent interactions, many of which are inherently multi-modal, i.e., they can lead to multiple distinct outcomes. To interact effectively, robots must recognize the possible…

机器人学 · 计算机科学 2025-08-12 Maulik Bhatt , Iman Askari , Yue Yu , Ufuk Topcu , Huazhen Fang , Negar Mehr

The marriage between mean-field theory and reinforcement learning has shown a great capacity to solve large-scale control problems with homogeneous agents. To break the homogeneity restriction of mean-field theory, a recent interest is to…

多智能体系统 · 计算机科学 2026-03-03 Yuanquan Hu , Xiaoli Wei , Junji Yan , Hengxi Zhang

We study the global convergence of policy optimization for finding the Nash equilibria (NE) in zero-sum linear quadratic (LQ) games. To this end, we first investigate the landscape of LQ games, viewing it as a nonconvex-nonconcave…

机器学习 · 计算机科学 2021-02-12 Kaiqing Zhang , Zhuoran Yang , Tamer Başar

Mean field games (MFGs) model interactions in large-population multi-agent systems through population distributions. Traditional learning methods for MFGs are based on fixed-point iteration (FPI), where policy updates and induced population…

机器学习 · 计算机科学 2025-02-17 Chenyu Zhang , Xu Chen , Xuan Di

In this paper, we consider a finite horizon, non-stationary, mean field games (MFG) with a large population of homogeneous players, sequentially making strategic decisions, where each player is affected by other players through an aggregate…

系统与控制 · 电气工程与系统科学 2020-04-07 Rajesh K Mishra , Deepanshu Vasal , Sriram Vishwanath

Mean field game (MFG) is an expressive modeling framework for systems with a continuum of interacting agents. While many approaches exist for solving the forward MFG, few have studied its \textit{inverse} problem. In this work, we seek to…

最优化与控制 · 数学 2025-07-28 Han Huang , Jiajia Yu , Tianyi Chen , Rongjie Lai

This paper develops a deep policy iteration method for high-dimensional finite-horizon mean-field games (MFG). We reformulate the game as a regenerative problem with deterministic cycles, which allows policy evaluation (PE), policy…

数值分析 · 数学 2026-05-18 Shuixin Fang , Shupeng Wang , Zhen Wu , Hui Zhang , Tao Zhou

This letter studies multi-agent reinforcement learning in partially observable Markov potential games. Solving this problem is challenging due to partial observability, decentralized information, and the curse of dimensionality. First, to…

多智能体系统 · 计算机科学 2026-04-02 Wonseok Yang , Thinh T. Doan