Related papers: Residual U-net with Self-Attention to Solve Multi-…
We study a collaborative multi-agent stochastic linear bandit setting, where $N$ agents that form a network communicate locally to minimize their overall regret. In this setting, each agent has its own linear bandit problem (its own reward…
The dynamic allocation problem, also known as the `multi-armed bandit' problem, simulates a situation in which an agent is faced with a tradeoff between actions that yield an immediate reward and actions whose benefits can only be perceived…
This paper studies an optimal stochastic impulse control problem in a finite horizon with a decision lag, by which we mean that after an impulse is made, a fixed number units of time has to be elapsed before the next impulse is allowed to…
In this paper, we study a stochastic recursive optimal control problem in which the objective functional is described by the solution of a backward stochastic differential equation driven by G-Brownian motion. Under standard assumptions, we…
This study investigates an optimal investment problem for an insurance company operating under the Cramer-Lundberg risk model, where investments are made in both a risky asset and a risk-free asset. In contrast to other literature that…
In this paper, a new yet indirect performance guaranteed framework is established to address the distributed tracking control problem for networked uncertain nonlinear strict-feedback systems with unknown time-varying gains under a directed…
We consider a class of exit time stochastic control problems for diffusion processes with discounted criterion, where the controller can utilize a given amount of resource, called "fuel". In contrast to the vast majority of existing…
In this report we derive the strategic (deterministic) allocation to bonds and stocks resulting in the optimal mean-variance trade-off on a given investment horizon. The underlying capital market features a mean-reverting process for equity…
In this paper, we explore a new class of stochastic control problems characterized by specific control constraints. Specifically, the admissible controls are subject to the ratcheting constraint, meaning they must be non-decreasing over…
We consider an optimal investment and consumption problem for a Black-Scholes financial market with stochastic coefficients driven by a diffusion process. We assume that an agent makes consumption and investment decisions based on CRRA…
In this paper, we investigate the distributed optimal control problem for a kind of nonlinear multi-agent systems. In particular,both the state and the system dynamic structures of each agent are private and can only be shared among…
This paper addresses the problem of distributed detection in multi-agent networks. Agents receive private signals about an unknown state of the world. The underlying state is globally identifiable, yet informative signals may be dispersed…
We consider the problem of maximizing portfolio value when an agent has a subjective view on asset value which differs from the traded market price. The agent's trades will have a price impact which affect the price at which the asset is…
Ad exchanges are widely used in platforms for online display advertising. Autonomous agents operating in these exchanges must learn policies for interacting profitably with a diverse, continually changing, but unknown market. We consider…
One of the most fundamental questions in quantitative finance is the existence of continuous-time diffusion models that fit market prices of a given set of options. Traditionally, one employs a mix of intuition, theoretical and empirical…
In this paper, we investigate whether deep reinforcement-learning agents interacting in a shared optimal-execution environment can sustain supra-competitive outcomes, in the sense of achieving lower implementation shortfalls than the…
We formulate a continuous-time competitive equilibrium model of irreversible capacity investment in which a continuum of heterogeneous producers supplies a single non-durable good subject to exogenous stochastic demand. Each producer…
We study the policy evaluation problem in multi-agent reinforcement learning where a group of agents, with jointly observed states and private local actions and rewards, collaborate to learn the value function of a given policy via local…
In this paper we consider a pairs trading financial market with the spread of risky assets defined by the Ornstein-Uhlenbeck (OU) process. We implement an optimal strategy for power utility functions for investment/consumption problem.…
We provide a unified approach to find equilibrium solutions for time-inconsistent problems with distribution dependent rewards, which are important to the study of behavioral finance and economics. Our approach is based on {\it equilibrium…