Related papers: An Actor-Critic Framework for Continuous-Time Jump…
We introduce a continuous policy-value iteration algorithm where the approximations of the value function of a stochastic control problem and the optimal control are simultaneously updated through Langevin-type dynamics. This framework…
Although actor-critic methods have been successful in practice, their theoretical analyses have several limitations. Specifically, existing theoretical work either sidesteps the exploration problem by making strong assumptions or analyzes…
We study equilibrium feedback strategies for a family of dynamic mean-variance problems with competition among a large group of agents. We assume that the time horizon is random and each agent's risk aversion depends dynamically on the…
In this paper, a Nash-type fictitious game framework is introduced to handle a time-inconsistent linear-quadratic optimal control. The Nash-type game in this framework is called fictitious as it is between the decision maker (called real…
Reinforcement learning, mathematically described by Markov Decision Problems, may be approached either through dynamic programming or policy search. Actor-critic algorithms combine the merits of both approaches by alternating between steps…
The finite state semi-Markov process is a generalization over the Markov chain in which the sojourn time distribution is any general distribution. In this article we provide a sufficient stochastic maximum principle for the optimal control…
Being able to fall safely is a necessary motor skill for humanoids performing highly dynamic tasks, such as running and jumping. We propose a new method to learn a policy that minimizes the maximal impulse during the fall. The optimization…
We study stochastic optimal control problems for (possibly degenerate) McKean-Vlasov controlled diffusions and obtain discrete-time as well as finite interacting particle approximations. (i) Under mild assumptions, we first prove the…
We develop an approach for solving time-consistent risk-sensitive stochastic optimization problems using model-free reinforcement learning (RL). Specifically, we assume agents assess the risk of a sequence of random variables using dynamic…
We establish the convergence of the deep actor-critic reinforcement learning algorithm presented in [Angiuli et al., 2023a] in the setting of continuous state and action spaces with an infinite discrete-time horizon. This algorithm provides…
Service platforms must determine rules for matching heterogeneous demand (customers) and supply (workers) that arrive randomly over time and may be lost if forced to wait too long for a match. Our objective is to maximize the cumulative…
This paper is concerned with the maximum principle and dynamic programming principle for mean-variance portfolio selection of jump diffusions and their relationship. First, the optimal portfolio and efficient frontier of the problem are…
In this paper, we consider the classic stochastic (dynamic) knapsack problem, a fundamental mathematical model in revenue management, with general time-varying random demand. Our main goal is to study the optimal policies, which can be…
Reinforcement learning in multi-agent scenarios is important for real-world applications but presents challenges beyond those seen in single-agent settings. We present an actor-critic algorithm that trains decentralized policies in…
Synchronizing decisions across multiple agents in realistic settings is problematic since it requires agents to wait for other agents to terminate and communicate about termination reliably. Ideally, agents should learn and execute…
Multi-agent reinforcement learning has been successfully applied to a number of challenging problems. Despite these empirical successes, theoretical understanding of different algorithms is lacking, primarily due to the curse of…
Building on recent advances in scientific machine learning and generative modeling for computational fluid dynamics, we propose a conditional score-based diffusion model designed for multi-scenarios fluid flow prediction. Our model…
In this paper, we investigate dynamic optimization problems featuring both stochastic control and optimal stopping in a finite time horizon. The paper aims to develop new methodologies, which are significantly different from those of mixed…
In this paper, we study a stochastic linear-quadratic control problem with random coefficients and regime switching on a horizon $[0,T\wedge\tau]$, where $\tau$ is a given random jump time for the underlying state process and $T$ is a…
Diffusion models have emerged as powerful tools for generative modeling, demonstrating exceptional capability in capturing target data distributions from large datasets. However, fine-tuning these massive models for specific downstream…