English
Related papers

Related papers: A Temporal Difference Method for Stochastic Contin…

200 papers

As autonomous systems become more ubiquitous in daily life, ensuring high performance with guaranteed safety is crucial. However, safety and performance could be competing objectives, which makes their co-optimization difficult.…

Robotics · Computer Science 2025-05-29 Manan Tayal , Aditya Singh , Shishir Kolathaya , Somil Bansal

In this note, we study a class of indefinite stochastic McKean-Vlasov linear-quadratic (LQ in short) control problem under the control taking nonnegative values. In contrast to the conventional issue, both the classical dynamic programming…

Optimization and Control · Mathematics 2023-10-05 Xun Li , Liangquan Zhang

This article is a continuation of a previous work where we studied infinite horizon control problems for which the dynamic, running cost and control space may be different in two half-spaces of some euclidian space $\R^N$. In this article…

Analysis of PDEs · Mathematics 2014-01-27 Guy Barles , Ariela Briani , Emmanuel Chasseigne

We study the problem of learning the optimal control policy for fine-tuning a given diffusion process, using general value function approximation. We develop a new class of algorithms by solving a variational inequality problem based on the…

Machine Learning · Computer Science 2025-09-03 Wenlong Mou

Time-consistency is an essential requirement in risk sensitive optimal control problems to make rational decisions. An optimization problem is time consistent if its solution policy does not depend on the time sequence of solving the…

Optimization and Control · Mathematics 2015-03-26 Yinlam Chow , Marco Pavone

In this paper, a new reinforcement learning (RL) method known as the method of temporal differential is introduced. Compared to the traditional temporal-difference learning method, it plays a crucial role in developing novel RL techniques…

Machine Learning · Computer Science 2020-06-02 Tao Bian , Zhong-Ping Jiang

This work provides a rigorous framework for studying continuous time control problems in uncertain environments. The framework considered models uncertainty in state dynamics as a measure on the space of functions. This measure is…

Optimization and Control · Mathematics 2018-02-22 Ryan Murray , Michele Palladino

The ergodic control problem for a non-degenerate controlled diffusion controlled through its drift is considered under a uniform stability condition that ensures the well-posedness of the associated Hamilton-Jacobi-Bellman (HJB) equation. A…

Optimization and Control · Mathematics 2019-03-20 Ari Arapostathis , Vivek S. Borkar

It is well known that time dependent Hamilton-Jacobi-Isaacs partial differential equations (HJ PDE), play an important role in analyzing continuous dynamic games and control theory problems. An important tool for such problems when they…

Optimization and Control · Mathematics 2016-05-09 Jérôme Darbon , Stanley Osher

We study an agent's lifecycle portfolio choice problem with stochastic labor income, borrowing constraints and a finite retirement date. Similarly to arXiv:2002.00201, wages evolve in a path-dependent way, but the presence of a finite…

Optimization and Control · Mathematics 2024-02-27 Sara Biagini , Enrico Biffis , Fausto Gozzi , Margherita Zanella

Considering that the decision-making environment faced by reinforcement learning (RL) agents is full of Knightian uncertainty, this paper describes the exploratory state dynamics equation in Knightian uncertainty to study the…

Optimization and Control · Mathematics 2026-01-27 Ziyu Li , Chen Fei , Weiyin Fei

We study the policy evaluation problem in multi-agent reinforcement learning where a group of agents, with jointly observed states and private local actions and rewards, collaborate to learn the value function of a given policy via local…

Optimization and Control · Mathematics 2021-11-08 Dongsheng Ding , Xiaohan Wei , Zhuoran Yang , Zhaoran Wang , Mihailo R. Jovanović

In this paper, we study the delayed stochastic recursive optimal control problem with a non-Lipschitz generator, in which both the dynamics of the control system and the recursive cost functional depend on the past path segment of the state…

Optimization and Control · Mathematics 2023-12-27 Jiaqiang Wen , Zhen Wu , Qi Zhang

We study the exploratory Hamilton--Jacobi--Bellman (HJB) equation arising from the entropy-regularized exploratory control problem, which was formulated by Wang, Zariphopoulou and Zhou (J. Mach. Learn. Res., 21, 2020) in the context of…

Optimization and Control · Mathematics 2021-09-22 Wenpin Tang , Paul Yuming Zhang , Xun Yu Zhou

In this paper, we study the optimal stopping problem in the so-called exploratory framework, in which the agent takes actions randomly conditioning on current state and an entropy-regularized term is added to the reward functional. Such a…

Optimization and Control · Mathematics 2023-09-04 Yuchao Dong

As autonomous robots move into complex, dynamic real-world environments, they must learn to navigate safely in real time, yet anticipating all possible behaviors is infeasible. We propose a composable, model-free reinforcement learning…

Robotics · Computer Science 2026-02-16 Xinhuan Sang , Abdelrahman Abdelgawad , Roberto Tron

In this paper, we present a scalable deep learning approach to solve opinion dynamics stochastic optimal control problems with mean field term coupling in the dynamics and cost function. Our approach relies on the probabilistic…

Multiagent Systems · Computer Science 2022-04-19 Tianrong Chen , Ziyi Wang , Evangelos A. Theodorou

Policy iteration (PI) is a recursive process of policy evaluation and improvement for solving an optimal decision-making/control problem, or in other words, a reinforcement learning (RL) problem. PI has also served as the fundamental for…

Artificial Intelligence · Computer Science 2021-04-06 Jaeyoung Lee , Richard S. Sutton

We extend the work on optimal investment and consumption of a population considered in [2] to a general stochastic setting over a finite time horizon. We incorporate the Cobb-Douglas production function in the capital dynamics while the…

Analysis of PDEs · Mathematics 2024-08-15 Hao Liu , Suresh P. Sethi , Tak Kwong Wong , Sheung Chi Phillip Yam

This paper develops a framework for establishing the existence of solutions to the equilibrium Hamilton-Jacobi-Bellman (EHJB) equation arising in time-inconsistent stochastic control problems. The time-inconsistency in our setting arises…

Optimization and Control · Mathematics 2026-04-07 Zhenhua Wang , Xiang Yu , Jingjie Zhang , Zhou Zhou