Related papers: An Actor-Critic Framework for Continuous-Time Jump…
Value-based algorithms are a cornerstone of off-policy reinforcement learning due to their simplicity and training stability. However, their use has traditionally been restricted to discrete action spaces, as they rely on estimating…
Flow-based generative models have shown remarkable success in text-to-image generation, yet fine-tuning them with intermediate feedback remains challenging, especially for continuous-time flow matching models. Most existing approaches…
Revisiting the continuous-time Mean-Variance (MV) Portfolio Optimization problem, we model the market dynamics with a jump-diffusion process and apply Reinforcement Learning (RL) techniques to facilitate informed exploration within the…
In this paper we prove a necessary condition of the optimal control problem for a class of general mean-field forward-backward stochastic systems with jumps in the case where the diffusion coefficients depend on control, the control set…
We develop an approach for two player constraint zero-sum and nonzero-sum stochastic differential games, which are modeled by Markov regime-switching jump-diffusion processes. We provide the relations between a usual stochastic optimal…
The option-critic architecture (Bacon, Harb, and Precup 2017) and several variants have successfully demonstrated the use of the options framework proposed by Sutton et al (Sutton, Precup, and Singh1999) to scale learning and planning in…
We study a family of mean field games with a state variable evolving as a multivariate jump diffusion process. The jump component is driven by a Poisson process with a time-dependent intensity function. All coefficients, i.e. drift,…
Deep Reinforcement Learning (DRL) algorithms for continuous action spaces are known to be brittle toward hyperparameters as well as \cut{being}sample inefficient. Soft Actor Critic (SAC) proposes an off-policy deep actor critic algorithm…
The stochastic $H_2/H_\infty$ control problem for continuous-time mean-field stochastic differential equations with Poisson jumps over finite horizon is investigated in this paper. Continuous and jump diffusion terms in the system depend…
In this paper, we consider a general time-inconsistent optimal control problem for a non homogeneous linear system, in which its state evolves according to a stochastic differential equation with deterministic coefficients, when the noise…
In this paper, we present a probability one convergence proof, under suitable conditions, of a certain class of actor-critic algorithms for finding approximate solutions to entropy-regularized MDPs using the machinery of stochastic…
Despite the empirical success of the actor-critic algorithm, its theoretical understanding lags behind. In a broader context, actor-critic can be viewed as an online alternating update algorithm for bilevel optimization, whose convergence…
In this paper, we propose a unified stochastic optimal control framework that integrates time-optimal control problems with classical stochastic optimal control formulations. Unlike conventional deterministic time-optimal control models,…
Recently, a theory for stochastic optimal control in non-linear dynamical systems in continuous space-time has been developed (Kappen, 2005). We apply this theory to collaborative multi-agent systems. The agents evolve according to a given…
We formulate and study a general time-varying multi-agent system where players repeatedly compete under incomplete information. Our work is motivated by scenarios commonly observed in online advertising and retail marketplaces, where agents…
Optimal processes in stochastic thermodynamics are a frontier for understanding the control and design of non-equilibrium systems, with broad practical applications in biology, chemistry, and nanoscale/mesoscale systems. Optimal mass…
In this paper, we investigate the infinite-horizon risk-constrained linear quadratic regulator problem (RC-QR), which augments the classical LQR formulation with a statistical constraint on the variability of the system state to incorporate…
We focus on a simulation-based optimization problem of choosing the best design from the feasible space. Although the simulation model can be queried with finite samples, its internal processing rule cannot be utilized in the optimization…
Achieving robust coordination and cooperation is a central challenge in multi-agent reinforcement learning (MARL). Uncovering the mechanisms underlying such emergent behaviors calls for a dynamical understanding of learn processes. In this…
In this paper, which is a continuation of the previously published discrete time paper we develop a theory for continuous time stochastic control problems which, in various ways, are time inconsistent in the sense that they do not admit a…