Related papers: SIMPOL Model for Solving Continuous-Time Heterogen…
In this paper, we present two effective policy learning algorithms for multi-agent online coordination(MA-OC) problem. The first one, \texttt{MA-SPL}, not only can achieve the optimal $(1-\frac{c}{e})$-approximation guarantee for the MA-OC…
To design effective digital interventions, experimenters face the challenge of learning decision policies that balance multiple objectives using offline data. Often, they aim to develop policies that maximize goal outcomes, while ensuring…
We study reinforcement learning in hybrid discrete-continuous action spaces, such as settings where the discrete component selects a regime (or index) and the continuous component optimizes within it -- a structure common in robotics,…
Multi-Agent Reinforcement Learning (MARL) -- where multiple agents learn to interact in a shared dynamic environment -- permeates across a wide range of critical applications. While there has been substantial progress on understanding the…
This paper studies continuous-time optimal contracting in a hierarchy problem which generalises the model of Sung (2015). The hierarchy is modeled by a series of interlinked principal-agent problems, leading to a sequence of Stackelberg…
A class of finite-state and discrete-time optimal control problems is introduced. The problems involve a large number of agents with independent dynamics, which interact through an aggregative term in the cost function. The problems are…
H{\infty} control of nonlinear continuous-time system depends on the solution of the Hamilton-Jacobi-Isaacs (HJI) equation, which has been proved impossible to obtain a closed-form solution due to the nonlinearity of HJI equation. In order…
In this paper, we propose a martingale-based neural network, SOC-MartNet, for solving high-dimensional Hamilton-Jacobi-Bellman (HJB) equations where no explicit expression is needed for the infimum of the Hamiltonian, $\inf_{u \in U}…
In this manuscript we consider optimal control problems of stochastic differential equations with delays in the state and in the control. First, we prove an equivalent Markovian reformulation on Hilbert spaces of the state equation. Then,…
The high-order complexity of human behaviour is likely the root cause of extreme difficulty in financial market projections. We consider that behavioural simulation can unveil systemic dynamics to support analysis. Simulating diverse human…
We consider an optimal investment and consumption problem for a Black-Scholes financial market with stochastic coefficients driven by a diffusion process. We assume that an agent makes consumption and investment decisions based on CRRA…
An optimal control problem is considered for a stochastic differential equation with the cost functional determined by a backward stochastic Volterra integral equation (BSVIE, for short). This kind of cost functional can cover the general…
This paper is concerned with the axiomatic foundation and explicit construction of a general class of optimality criteria that can be used for investment problems with multiple time horizons, or when the time horizon is not known in…
This paper presents novel solutions of the data-based synchronization problem for continuous-time multiagent systems. We consider the cases of homogeneous and heterogeneous systems. First, we obtain a data-based representation of the…
In this paper, we focus on the problem of optimal portfolio-consumption policies in a multi-asset financial market, where the n risky assets follow Exponential Ornstein-Uhlenbeck processes, along with one risk-free bond. The investor's…
The paper studies a system of first order Hamilton-Jacobi equations with discontinuous coefficients, arising from a model of deterministic optimal debt management in infinite time horizon, with exponential discount and currency devaluation.…
We propose a scalable, policy-centric framework for continuous-time multi-asset portfolio-consumption optimization under inequality constraints. Our method integrates neural policies with Pontryagin's Maximum Principle (PMP) and enforces…
We consider a class of exit time stochastic control problems for diffusion processes with discounted criterion, where the controller can utilize a given amount of resource, called "fuel". In contrast to the vast majority of existing…
We propose a new policy gradient method, named homotopic policy mirror descent (HPMD), for solving discounted, infinite horizon MDPs with finite state and action spaces. HPMD performs a mirror descent type policy update with an additional…
We propose a scalable method for computing global solutions of nonlinear, high-dimensional dynamic stochastic economic models. First, within a time iteration framework, we approximate economic policy functions using an adaptive,…