Related papers: Policy Gradient for Continuous-Time Mean-Field Con…
We approach the continuous-time mean-variance (MV) portfolio selection with reinforcement learning (RL). The problem is to achieve the best tradeoff between exploration and exploitation, and is formulated as an entropy-regularized, relaxed…
This note re-visits the rolling-horizon control approach to the problem of a Markov decision process (MDP) with infinite-horizon discounted expected reward criterion. Distinguished from the classical value-iteration approach, we develop an…
We consider the problem of non-stationary reinforcement learning (RL) in the infinite-horizon average-reward setting. We model it by a Markov Decision Process with time-varying rewards and transition probabilities, with a variation budget…
In computational reinforcement learning, a growing body of work seeks to express an agent's model of the world through predictions about future sensations. In this manuscript we focus on predictions expressed as General Value Functions:…
In this paper, we introduce discrete-time linear mean-field games subject to an infinite-horizon discounted-cost optimality criterion. The state space of a generic agent is a compact Borel space. At every time, each agent is randomly…
Large-scale competitive platforms are interacting multi-agent systems in which latent skills drift over time and pairwise interactions are shaped by matchmaking. We study a controlled rating dynamics in the mean-field limit and derive a…
In this paper, we investigate a model-free optimal control design that minimizes an infinite horizon average expected quadratic cost of states and control actions subject to a probabilistic risk or chance constraint using input-output data.…
We consider the static output feedback control for Linear Quadratic Regulator problems with structured constraints under the assumption that system parameters are unknown. To solve the problem in the model free setting, we propose the…
For over a decade, model-based reinforcement learning has been seen as a way to leverage control-based domain knowledge to improve the sample-efficiency of reinforcement learning agents. While model-based agents are conceptually appealing,…
This paper presents a novel form of policy gradient for model-free reinforcement learning (RL) with improved exploration properties. Current policy-based methods use entropy regularization to encourage undirected exploration of the reward…
This paper studies policy learning for continuous treatments from observational data. Continuous treatments present more significant challenges than discrete ones because population welfare may need nonparametric estimation, and policy…
This paper studies dynamic mean-variance (MV) asset allocation problems in general incomplete markets. Besides of the conventional MV objective on portfolio's terminal wealth, our framework can accommodate running MV objectives with general…
Standard reinforcement learning methods aim to master one way of solving a task whereas there may exist multiple near-optimal policies. Being able to identify this collection of near-optimal policies can allow a domain expert to efficiently…
We introduce a novel extension to robust control theory that explicitly addresses uncertainty in the value function's gradient, a form of uncertainty endemic to applications like reinforcement learning where value functions are…
This paper is devoted to a class of finite horizon deterministic mean field games with Grushin type dynamics, state constraints and nonlocal coupling. First, we consider the optimal control problem that each agent aims to solve when the…
We consider an infinite horizon discounted optimal control problem for piecewise deterministic Markov processes, where a piecewise open-loop control acts continuously on the jump dynamics and on the deterministic flow. For this class of…
Reinforcement learning algorithms such as the deep deterministic policy gradient algorithm (DDPG) has been widely used in continuous control tasks. However, the model-free DDPG algorithm suffers from high sample complexity. In this paper we…
We study the problem of parameter estimation for large exchangeable interacting particle systems when a sample of discrete observations from a single particle is known. We propose a novel method based on martingale estimating functions…
Policy gradient algorithms have driven many recent advancements in language model reasoning. An appealing property is their ability to learn from exploration on their own trajectories, a process crucial for fostering diverse and creative…
We study the large-population limit of interacting particle systems evolving on adaptive dynamical networks, motivated in particular by models of opinion dynamics. In such systems, agents interact through weighted graphs whose structure…