English
Related papers

Related papers: Continuous-time q-learning for mean-field control …

200 papers

This paper investigates an indefinite linear-quadratic partially observed mean-field game with common noise, incorporating both state-average and control-average effects. In our model, each agent's state is observed through both individual…

Optimization and Control · Mathematics 2025-08-05 Tian Chen , Tianyang Nie , Zhen Wu

Considering that the decision-making environment faced by reinforcement learning (RL) agents is full of Knightian uncertainty, this paper describes the exploratory state dynamics equation in Knightian uncertainty to study the…

Optimization and Control · Mathematics 2026-01-27 Ziyu Li , Chen Fei , Weiyin Fei

This paper presents a novel model-free and fully data-driven policy iteration scheme for quadratic regulation of linear dynamics with state- and input-multiplicative noise. The implementation is similar to the least-squares temporal…

Optimization and Control · Mathematics 2022-12-05 Peter Coppens , Panagiotis Patrinos

Regularized Markov Decision Processes serve as models of sequential decision making under uncertainty wherein the decision maker has limited information processing capacity and/or aversion to model ambiguity. With functional approximation,…

Artificial Intelligence · Computer Science 2025-02-11 Jiachen Xi , Alfredo Garcia , Petar Momcilovic

This paper studies the continuous-time reinforcement learning in jump-diffusion models by featuring the q-learning (the continuous-time counterpart of Q-learning) under Tsallis entropy regularization. Contrary to the Shannon entropy, the…

Optimization and Control · Mathematics 2026-02-16 Lijun Bo , Yijie Huang , Xiang Yu , Tingting Zhang

We study a decentralized version of Moving Agents in Formation (MAiF), a variant of Multi-Agent Path Finding aiming to plan collision-free paths for multiple agents with the dual objectives of reaching their goals quickly while maintaining…

Robotics · Computer Science 2024-10-17 Qiushi Lin , Hang Ma

We study the discrete-time linear-quadratic (LQ) control model using reinforcement learning (RL). Using entropy to measure the cost of exploration, we prove that the optimal feedback policy for the problem must be Gaussian type. Then, we…

Machine Learning · Statistics 2025-02-05 Lucky Li

We revisit the unified two-timescale Q-learning algorithm as initially introduced by Angiuli et al. \cite{angiuli2022unified}. This algorithm demonstrates efficacy in solving mean field game (MFG) and mean field control (MFC) problems,…

Optimization and Control · Mathematics 2024-05-30 Jing An , Jianfeng Lu , Yue Wu , Yang Xiang

We investigate reinforcement learning in the setting of Markov decision processes for a large number of exchangeable agents interacting in a mean field manner. Applications include, for example, the control of a large number of robots…

Optimization and Control · Mathematics 2025-04-30 René Carmona , Mathieu Laurière , Zongjun Tan

This paper employs a policy iteration reinforcement learning (RL) method to study continuous-time linear-quadratic mean-field control problems in infinite horizon. The drift and diffusion terms in the dynamics involve the states, the…

Optimization and Control · Mathematics 2024-11-05 Na Li , Xun Li , Zuo Quan Xu

Reinforcement learning is a powerful tool to learn the optimal policy of possibly multiple agents by interacting with the environment. As the number of agents grow to be very large, the system can be approximated by a mean-field problem.…

Optimization and Control · Mathematics 2020-08-18 Weichen Wang , Jiequn Han , Zhuoran Yang , Zhaoran Wang

In deep reinforcement learning, estimating the value function to evaluate the quality of states and actions is essential. The value function is often trained using the least squares method, which implicitly assumes a Gaussian error…

Machine Learning · Computer Science 2024-03-28 Motoki Omura , Takayuki Osa , Yusuke Mukuta , Tatsuya Harada

This paper is a continuation of Ishitani and Kato (2015), in which we derived a continuous-time value function corresponding to an optimal execution problem with uncertain market impact as the limit of a discrete-time value function. Here,…

Trading and Market Microstructure · Quantitative Finance 2015-11-10 Kensuke Ishitani , Takashi Kato

We propose a deep learning approach to compute mean field control problems with individual noises. The problem consists of the Fokker-Planck (FP) equation and the Hamilton-Jacobi-Bellman (HJB) equation. Using the differential of the…

Optimization and Control · Mathematics 2025-05-20 Mo Zhou , Stanley Osher , Wuchen Li

We study the problem of optimal portfolio selection under stochastic volatility within a continuous time reinforcement learning framework with portfolio constraints. Exploration is modeled through entropy-regularized relaxed controls, where…

Mathematical Finance · Quantitative Finance 2026-04-27 Thai Nguyen , Pertiny Nkuize

In the past few years, off-policy reinforcement learning methods have shown promising results in their application for robot control. Deep Q-learning, however, still suffers from poor data-efficiency and is susceptible to stochasticity in…

Machine Learning · Computer Science 2020-08-17 Gabriel Kalweit , Maria Huegle , Joschka Boedecker

The exploration-exploitation dilemma has been a central challenge in reinforcement learning (RL) with complex model classes. In this paper, we propose a new algorithm, Monotonic Q-Learning with Upper Confidence Bound (MQL-UCB) for RL with…

Machine Learning · Computer Science 2025-10-06 Heyang Zhao , Jiafan He , Quanquan Gu

Continuous-time reinforcement learning offers an appealing formalism for describing control problems in which the passage of time is not naturally divided into discrete increments. Here we consider the problem of predicting the distribution…

Machine Learning · Computer Science 2022-06-20 Harley Wiltzer , David Meger , Marc G. Bellemare

We study policy gradient for mean-field control in continuous time in a reinforcement learning setting. By considering randomised policies with entropy regularisation, we derive a gradient expectation representation of the value function,…

Machine Learning · Statistics 2023-03-14 Noufel Frikha , Maximilien Germain , Mathieu Laurière , Huyên Pham , Xuanye Song

In this paper, two Q-learning (QL) methods are proposed and their convergence theories are established for addressing the model-free optimal control problem of general nonlinear continuous-time systems. By introducing the Q-function for…

Systems and Control · Computer Science 2014-10-14 Biao Luo , Derong Liu , Tingwen Huang