English
Related papers

Related papers: An Actor-Critic Framework for Continuous-Time Jump…

200 papers

We introduce a continuous policy-value iteration algorithm where the approximations of the value function of a stochastic control problem and the optimal control are simultaneously updated through Langevin-type dynamics. This framework…

Optimization and Control · Mathematics 2025-06-11 Qi Feng , Gu Wang

Although actor-critic methods have been successful in practice, their theoretical analyses have several limitations. Specifically, existing theoretical work either sidesteps the exploration problem by making strong assumptions or analyzes…

Machine Learning · Computer Science 2026-04-02 Max Qiushi Lin , Reza Asad , Kevin Tan , Haque Ishfaq , Csaba Szepesvari , Sharan Vaswani

We study equilibrium feedback strategies for a family of dynamic mean-variance problems with competition among a large group of agents. We assume that the time horizon is random and each agent's risk aversion depends dynamically on the…

Optimization and Control · Mathematics 2026-05-05 Xiaoqing Liang , Jie Xiong , Ying Yang

In this paper, a Nash-type fictitious game framework is introduced to handle a time-inconsistent linear-quadratic optimal control. The Nash-type game in this framework is called fictitious as it is between the decision maker (called real…

Optimization and Control · Mathematics 2021-10-04 Yuan-Hua Ni , Binbin Si , Xinzhen Zhang

Reinforcement learning, mathematically described by Markov Decision Problems, may be approached either through dynamic programming or policy search. Actor-critic algorithms combine the merits of both approaches by alternating between steps…

Machine Learning · Computer Science 2023-01-31 Harshat Kumar , Alec Koppel , Alejandro Ribeiro

The finite state semi-Markov process is a generalization over the Markov chain in which the sojourn time distribution is any general distribution. In this article we provide a sufficient stochastic maximum principle for the optimal control…

Optimization and Control · Mathematics 2014-07-14 Amogh Deshpande

Being able to fall safely is a necessary motor skill for humanoids performing highly dynamic tasks, such as running and jumping. We propose a new method to learn a policy that minimizes the maximal impulse during the fall. The optimization…

Robotics · Computer Science 2017-04-21 Visak CV Kumar , Sehoon Ha , C Karen Liu

We study stochastic optimal control problems for (possibly degenerate) McKean-Vlasov controlled diffusions and obtain discrete-time as well as finite interacting particle approximations. (i) Under mild assumptions, we first prove the…

Optimization and Control · Mathematics 2025-10-27 Somnath Pradhan , Serdar Yuksel

We develop an approach for solving time-consistent risk-sensitive stochastic optimization problems using model-free reinforcement learning (RL). Specifically, we assume agents assess the risk of a sequence of random variables using dynamic…

Machine Learning · Computer Science 2022-12-01 Anthony Coache , Sebastian Jaimungal

We establish the convergence of the deep actor-critic reinforcement learning algorithm presented in [Angiuli et al., 2023a] in the setting of continuous state and action spaces with an infinite discrete-time horizon. This algorithm provides…

Optimization and Control · Mathematics 2025-11-11 Jean-Pierre Fouque , Mathieu Laurière , Mengrui Zhang

Service platforms must determine rules for matching heterogeneous demand (customers) and supply (workers) that arrive randomly over time and may be lost if forced to wait too long for a match. Our objective is to maximize the cumulative…

Optimization and Control · Mathematics 2023-12-19 Angelos Aveklouris , Levi DeValve , Maximiliano Stock , Amy R. Ward

This paper is concerned with the maximum principle and dynamic programming principle for mean-variance portfolio selection of jump diffusions and their relationship. First, the optimal portfolio and efficient frontier of the problem are…

Portfolio Management · Quantitative Finance 2025-08-05 Qiyue Zhang , Jingtao Shi

In this paper, we consider the classic stochastic (dynamic) knapsack problem, a fundamental mathematical model in revenue management, with general time-varying random demand. Our main goal is to study the optimal policies, which can be…

Optimization and Control · Mathematics 2018-07-19 Yingdong Lu

Reinforcement learning in multi-agent scenarios is important for real-world applications but presents challenges beyond those seen in single-agent settings. We present an actor-critic algorithm that trains decentralized policies in…

Machine Learning · Computer Science 2019-05-29 Shariq Iqbal , Fei Sha

Synchronizing decisions across multiple agents in realistic settings is problematic since it requires agents to wait for other agents to terminate and communicate about termination reliably. Ideally, agents should learn and execute…

Machine Learning · Computer Science 2022-10-12 Yuchen Xiao , Weihao Tan , Christopher Amato

Multi-agent reinforcement learning has been successfully applied to a number of challenging problems. Despite these empirical successes, theoretical understanding of different algorithms is lacking, primarily due to the curse of…

Machine Learning · Computer Science 2021-12-28 Yuwei Luo , Zhuoran Yang , Zhaoran Wang , Mladen Kolar

Building on recent advances in scientific machine learning and generative modeling for computational fluid dynamics, we propose a conditional score-based diffusion model designed for multi-scenarios fluid flow prediction. Our model…

Machine Learning · Computer Science 2025-06-02 Wilfried Genuist , Éric Savin , Filippo Gatti , Didier Clouteau

In this paper, we investigate dynamic optimization problems featuring both stochastic control and optimal stopping in a finite time horizon. The paper aims to develop new methodologies, which are significantly different from those of mixed…

Portfolio Management · Quantitative Finance 2014-06-27 Xiongfei Jian , Xun Li , Fahuai Yi

In this paper, we study a stochastic linear-quadratic control problem with random coefficients and regime switching on a horizon $[0,T\wedge\tau]$, where $\tau$ is a given random jump time for the underlying state process and $T$ is a…

Optimization and Control · Mathematics 2022-01-19 Ying Hu , Xiaomin Shi , Zuo Quan Xu

Diffusion models have emerged as powerful tools for generative modeling, demonstrating exceptional capability in capturing target data distributions from large datasets. However, fine-tuning these massive models for specific downstream…

Machine Learning · Computer Science 2025-09-01 Yinbin Han , Meisam Razaviyayn , Renyuan Xu
‹ Prev 1 3 4 5 6 7 10 Next ›