English
Related papers

Related papers: Unified continuous-time q-learning for mean-field …

200 papers

This paper establishes a data-driven solution for infinite horizon linear quadratic Gaussian Mean Field Games with network-coupled heterogeneous agent populations where the dynamics of the agents are unknown. The solution technique relies…

Systems and Control · Electrical Eng. & Systems 2026-02-17 Jean Zhu , Shuang Gao

This paper addresses zero-sum ``turn'' games, in which only one player can make decisions at each state. We show that pure saddle-point state-feedback policies for turn games can be constructed from dynamic programming fixed-point equations…

Systems and Control · Electrical Eng. & Systems 2025-09-18 Sean Anderson , Chris Darken , João Hespanha

We present an algorithm for learning an approximate action-value soft Q-function in the relative entropy regularised reinforcement learning setting, for which an optimal improved policy can be recovered in closed form. We use recent…

Machine Learning · Computer Science 2019-11-06 Jonas Degrave , Abbas Abdolmaleki , Jost Tobias Springenberg , Nicolas Heess , Martin Riedmiller

Multi-agent reinforcement learning (MARL) achieves significant empirical successes. However, MARL suffers from the curse of many agents. In this paper, we exploit the symmetry of agents in MARL. In the most generic form, we study a…

Machine Learning · Computer Science 2020-06-23 Lingxiao Wang , Zhuoran Yang , Zhaoran Wang

We establish a probabilistic framework for analysing extended mean-field games with multi-dimensional singular controls and state-dependent jump dynamics and costs. Two key challenges arise when analysing such games: the state dynamics may…

Optimization and Control · Mathematics 2024-11-25 Robert Denkert , Ulrich Horst

We study a family of mean field games with a state variable evolving as a multivariate jump diffusion process. The jump component is driven by a Poisson process with a time-dependent intensity function. All coefficients, i.e. drift,…

Probability · Mathematics 2020-07-14 Chiara Benazzoli , Luciano Campi , Luca Di Persio

This paper studies a class of partial information linear-quadratic mean-field game problems. A general stochastic large-population system is considered, where the diffusion term of the dynamic of each agent can depend on the state and…

Optimization and Control · Mathematics 2022-03-22 Min Li , Tianyang Nie , Zhen Wu

Off-Policy reinforcement learning (RL) is an important class of methods for many problem domains, such as robotics, where the cost of collecting data is high and on-policy methods are consequently intractable. Standard methods for applying…

Artificial Intelligence · Computer Science 2019-07-03 Riley Simmons-Edler , Ben Eisner , Eric Mitchell , Sebastian Seung , Daniel Lee

Modeling joint probability distributions is an important task in a wide variety of fields. One popular technique for this employs a family of multivariate distributions with uniform marginals called copulas. While the theory of modeling…

Federated learning enables a collaborative training and optimization of global models among a group of devices without sharing local data samples. However, the heterogeneity of data in federated learning can lead to unfair representation of…

Machine Learning · Computer Science 2023-11-03 Weikang Chen , Junping Du , Yingxia Shao , Jia Wang , Yangxi Zhou

Mean-Field Control (MFC) has recently been proven to be a scalable tool to approximately solve large-scale multi-agent reinforcement learning (MARL) problems. However, these studies are typically limited to unconstrained cumulative reward…

Machine Learning · Computer Science 2024-09-11 Washim Uddin Mondal , Vaneet Aggarwal , Satish V. Ukkusuri

Previous work in hierarchical reinforcement learning has faced a dilemma: either ignore the values of different possible exit states from a subroutine, thereby risking suboptimal behavior, or represent those values explicitly thereby…

Machine Learning · Computer Science 2012-07-02 Bhaskara Marthi , Stuart Russell , David Andre

Q-functions are widely used in discrete-time learning and control to model future costs arising from a given control policy, when the initial state and input are given. Although some of their properties are understood, Q-functions…

Optimization and Control · Mathematics 2019-02-21 Joseph Warrington

In the Economic Nonlinear Model Predictive (ENMPC) context, closed-loop stability relates to the existence of a storage function satisfying a dissipation inequality. Finding the storage function is in general -- for nonlinear dynamics and…

Systems and Control · Electrical Eng. & Systems 2021-10-26 Arash Bahari Kordabad , Sebastien Gros

In this work, we study a class of mean-field linear quadratic Gaussian (LQG) problems. Under suitable conditions, explicit solutions of the distribution-dependent optimal control problems are obtained. Riccati systems are derived by…

Probability · Mathematics 2020-08-28 Yun Li , Qingshuo Song , Fuke Wu , George Yin

Revisiting the continuous-time Mean-Variance (MV) Portfolio Optimization problem, we model the market dynamics with a jump-diffusion process and apply Reinforcement Learning (RL) techniques to facilitate informed exploration within the…

Portfolio Management · Quantitative Finance 2025-12-11 Yuling Max Chen , Bin Li , David Saunders

In this article, we apply a probabilistic approach to study general mean field type control (MFTC) problems with jump-diffusions, and give the first global-in-time solution. We allow the drift coefficient $b$ and the diffusion coefficient…

Probability · Mathematics 2025-10-01 Alain Bensoussan , Ziyu Huang , Shanjian Tang , Sheung Chi Phillip Yam

In many real-world settings, a team of agents must coordinate their behaviour while acting in a decentralised way. At the same time, it is often possible to train the agents in a centralised fashion in a simulated or laboratory setting,…

We study a model-free federated linear quadratic regulator (LQR) problem where M agents with unknown, distinct yet similar dynamics collaboratively learn an optimal policy to minimize an average quadratic cost while keeping their data…

Optimization and Control · Mathematics 2023-08-24 Han Wang , Leonardo F. Toso , Aritra Mitra , James Anderson

The objective of meta-learning is to exploit the knowledge obtained from observed tasks to improve adaptation to unseen tasks. As such, meta-learners are able to generalize better when they are trained with a larger number of observed tasks…

Machine Learning · Computer Science 2022-10-11 Mert Kayaalp , Stefan Vlaski , Ali H. Sayed