中文
相关论文

相关论文: Recurrent Structural Policy Gradient for Partially…

200 篇论文

Tractable yet expressive density estimators are a key building block of probabilistic machine learning. While sum-product networks (SPNs) offer attractive inference capabilities, obtaining structures large enough to fit complex,…

机器学习 · 计算机科学 2019-08-12 Fabrizio Ventola , Karl Stelzner , Alejandro Molina , Kristian Kersting

Recent advances in mean-field game literature enable the reduction of large-scale multi-agent problems to tractable interactions between a representative agent and a population distribution. However, existing approaches typically assume a…

多智能体系统 · 计算机科学 2026-02-17 Bhavini Jeloka , Yue Guan , Panagiotis Tsiotras

Mean field games (MFG) and mean field control (MFC) problems have been introduced to study large populations of strategic players. They correspond respectively to non-cooperative or cooperative scenarios, where the aim is to find the Nash…

计算机科学与博弈论 · 计算机科学 2023-12-19 Rene Carmona , Gokce Dayanikli , Francois Delarue , Mathieu Lauriere

We introduce the receding-horizon policy gradient (RHPG) algorithm, the first PG algorithm with provable global convergence in learning the optimal linear estimator designs, i.e., the Kalman filter (KF). Notably, the RHPG algorithm does not…

最优化与控制 · 数学 2023-09-12 Xiangyuan Zhang , Saviz Mowlavi , Mouhacine Benosman , Tamer Başar

We consider first order variational MFG in the whole space, with aggregative interactions and density constraints, such that the stationary states of the game are contained in two isolated compact sets of mass distributions with finite…

偏微分方程分析 · 数学 2020-12-15 Annalisa Cesaroni , Marco Cirant

Even when confronted with the same data, agents often disagree on a model of the real-world. Here, we address the question of how interacting heterogenous agents, who disagree on what model the real-world follows, optimize their trading…

数理金融 · 定量金融 2019-12-13 Philippe Casgrain , Sebastian Jaimungal

We examine global non-asymptotic convergence properties of policy gradient methods for multi-agent reinforcement learning (RL) problems in Markov potential games (MPG). To learn a Nash equilibrium of an MPG in which the size of state space…

机器学习 · 计算机科学 2022-08-08 Dongsheng Ding , Chen-Yu Wei , Kaiqing Zhang , Mihailo R. Jovanović

This paper investigates the simultaneous reconstruction of the running cost function and the internal topological structure within the mean-field games (MFG) system utilizing partial boundary data. The inverse problem is notably challenging…

最优化与控制 · 数学 2024-08-20 Ming-Hui Ding , Hongyu Liu , Guang-Hui Zheng

Multi-agent deep reinforcement learning makes optimal decisions dependent on system states observed by agents, but any uncertainty on the observations may mislead agents to take wrong actions. The Mean-Field Actor-Critic reinforcement…

机器学习 · 计算机科学 2023-06-01 Ziyuan Zhou , Guanjun Liu

We explore a Federated Reinforcement Learning (FRL) problem where $N$ agents collaboratively learn a common policy without sharing their trajectory data. To date, existing FRL work has primarily focused on agents operating in the same or…

机器学习 · 计算机科学 2024-06-03 Han Wang , Sihong He , Zhili Zhang , Fei Miao , James Anderson

Online off-policy reinforcement learning (RL) is shaped by two coupled choices: the policy class and the update rule. Gaussian policies are fast and have tractable entropy, but struggle with multimodal action distributions. Generative…

机器学习 · 计算机科学 2026-05-22 Zeyuan Wang , Da Li , Yulin Chen , Yuehu Gong , Yanming Guo , Ye Shi , Liang Bai , Tianyuan Yu , Yanwei Fu

The mean field games (MFG) theory has broad application in mathematical modeling of social phenomena. The Mean Field Games System (MFGS) is the key to the MFG theory. This is a system of two nonlinear parabolic partial differential…

偏微分方程分析 · 数学 2024-02-26 Michael V. Klibanov , Jingzhi Li , Hongyu Liu

An iterative finite difference scheme for mean field games (MFGs) is proposed. The target MFGs are derived from control problems for multidimensional systems with advection terms. For such MFGs, linearization using the Cole-Hopf…

最优化与控制 · 数学 2023-04-26 Daisuke Inoue , Yuji Ito , Takahito Kashiwabara , Norikazu Saito , Hiroaki Yoshida

A key challenge in multi-agent systems is the design of intelligent agents solving real-world tasks in close interaction with other agents (e.g. humans), thereby being confronted with a variety of behavioral variations and limited knowledge…

多智能体系统 · 计算机科学 2020-07-13 Julian Bernhard , Alois Knoll

We propose a mean field game (MFG) framework to model the evolution of renewable energy production in competitive electricity markets. Producers interact through the spot price while optimising their profits under production, installation,…

最优化与控制 · 数学 2026-03-25 Luciano Campi , Zhuoshu Wu

Recent advances in deep learning has witnessed many innovative frameworks that solve high dimensional mean-field games (MFG) accurately and efficiently. These methods, however, are restricted to solving single-instance MFG and demands…

机器学习 · 计算机科学 2024-04-25 Han Huang , Rongjie Lai

Sample inefficiency is a long-lasting problem in reinforcement learning (RL). The state-of-the-art estimates the optimal action values while it usually involves an extensive search over the state-action space and unstable optimization.…

机器学习 · 计算机科学 2019-11-27 Kaixiang Lin , Jiayu Zhou

We develop the theory of linear-quadratic (LQ) mean field games (MFGs) in Hilbert spaces with common noise modeled by an infinite-dimensional Wiener process that affects the dynamics of all agents. In the presence of common noise, the…

最优化与控制 · 数学 2026-05-28 Hanchao Liu , Dena Firoozi

Policy gradient methods have been successfully applied to many complex reinforcement learning problems. However, policy gradient methods suffer from high variance, slow convergence, and inefficient exploration. In this work, we introduce a…

机器学习 · 计算机科学 2017-04-11 Yang Liu , Prajit Ramachandran , Qiang Liu , Jian Peng

The theory of continuous-time reinforcement learning (RL) has progressed rapidly in recent years. While the ultimate objective of RL is typically to learn deterministic control policies, most existing continuous-time RL methods rely on…

机器学习 · 计算机科学 2026-03-17 Ziheng Cheng , Xin Guo , Yufei Zhang