English
Related papers

Related papers: A Temporal Difference Method for Stochastic Contin…

200 papers

Linear Temporal Logic (LTL) is widely used to specify high-level objectives for system policies, and it is highly desirable for autonomous systems to learn the optimal policy with respect to such specifications. However, learning the…

Machine Learning · Computer Science 2023-10-26 Daqian Shao , Marta Kwiatkowska

We consider the problem of Reinforcement Learning for nonlinear stochastic dynamical systems. We show that in the RL setting, there is an inherent ``Curse of Variance" in addition to Bellman's infamous ``Curse of Dimensionality", in…

Machine Learning · Computer Science 2021-07-30 Raman Goyal , Suman Chakravorty , Ran Wang , Mohamed Naveed Gul Mohamed

In this paper, we investigate a sparse optimal control of continuous-time stochastic systems. We adopt the dynamic programming approach and analyze the optimal control via the value function. Due to the non-smoothness of the $L^0$ cost…

Optimization and Control · Mathematics 2021-09-17 Kaito Ito , Takuya Ikeda , Kenji Kashima

This paper studies the stochastic optimal control of jump-diffusion processes and the associated fully nonlinear backward stochastic Hamilton--Jacobi--Bellman (BSHJB) equations. We establish the dynamic programming principle (DPP) via…

Optimization and Control · Mathematics 2026-05-21 Dunxiang Liang , Qingxin Meng

A gradient-enhanced functional tensor train cross approximation method for the resolution of the Hamilton-Jacobi-Bellman (HJB) equations associated to optimal feedback control of nonlinear dynamics is presented. The procedure uses samples…

Numerical Analysis · Mathematics 2023-02-23 Sergey Dolgov , Dante Kalise , Luca Saluzzi

Gradient-regularized value learning methods improve sample efficiency by leveraging learned models of transition dynamics and rewards to estimate return gradients. However, existing approaches, such as MAGE, struggle in stochastic or noisy…

Machine Learning · Computer Science 2026-03-04 Baptiste Debes , Tinne Tuytelaars

Optimal control of general nonlinear systems is a central challenge in automation. Enabled by powerful function approximators, data-driven approaches to control have recently successfully tackled challenging applications. However, such…

Systems and Control · Electrical Eng. & Systems 2023-06-21 Hany Abdulsamad , Jan Peters

Standard reinforcement learning (RL) aims to find an optimal policy that identifies the best action for each state. However, in healthcare settings, many actions may be near-equivalent with respect to the reward (e.g., survival). We…

Machine Learning · Computer Science 2020-07-27 Shengpu Tang , Aditya Modi , Michael W. Sjoding , Jenna Wiens

We describe a new approach for managing aleatoric uncertainty in the Reinforcement Learning (RL) paradigm. Instead of selecting actions according to a single statistic, we propose a distributional method based on the second-order stochastic…

Machine Learning · Computer Science 2020-10-08 John D. Martin , Michal Lyskawinski , Xiaohu Li , Brendan Englot

This paper addresses the model-free nonlinear optimal problem with generalized cost functional, and a data-based reinforcement learning technique is developed. It is known that the nonlinear optimal control problem relies on the solution of…

Systems and Control · Computer Science 2013-11-20 Biao Luo , Huai-Ning Wu , Tingwen Huang , Derong Liu

Autonomous systems have witnessed a rapid increase in their capabilities, but it remains a challenge for them to perform tasks both effectively and safely. The fact that performance and safety can sometimes be competing objectives renders…

Systems and Control · Electrical Eng. & Systems 2024-12-04 Hao Wang , Adityaya Dhande , Somil Bansal

In this paper we investigate a dynamic stochastic portfolio optimization problem involving both the expected terminal utility and intertemporal utility maximization. We solve the problem by means of a solution to a fully nonlinear…

Portfolio Management · Quantitative Finance 2019-03-26 Sona Kilianova , Daniel Sevcovic

In this manuscript we consider a class optimal control problem for stochastic differential delay equations. First, we rewrite the problem in a suitable infinite-dimensional Hilbert space. Then, using the dynamic programming approach, we…

Optimization and Control · Mathematics 2023-02-20 Filippo de Feo , Salvatore Federico , Andrzej Święch

This is the first in a series of papers in which we study an efficient approximation scheme for solving the Hamilton-Jacobi-Bellman equation for multi-dimensional problems in stochastic control theory. The method is a combination of a WKB…

Computational Finance · Quantitative Finance 2014-06-26 Sakda Chaiworawitkul , Patrick S. Hagan , Andrew Lesniewski

We develop an approach for solving time-consistent risk-sensitive stochastic optimization problems using model-free reinforcement learning (RL). Specifically, we assume agents assess the risk of a sequence of random variables using dynamic…

Machine Learning · Computer Science 2022-12-01 Anthony Coache , Sebastian Jaimungal

This paper studies reinforcement learning (RL) in doubly inhomogeneous environments under temporal non-stationarity and subject heterogeneity. In a number of applications, it is commonplace to encounter datasets generated by system dynamics…

Machine Learning · Statistics 2025-03-18 Liyuan Hu , Mengbing Li , Chengchun Shi , Zhenke Wu , Piotr Fryzlewicz

We mathematically analyze and numerically study an actor-critic machine learning algorithm for solving high-dimensional Hamilton-Jacobi-Bellman (HJB) partial differential equations from stochastic control theory. The architecture of the…

Optimization and Control · Mathematics 2026-05-20 Samuel N. Cohen , Jackson Hebner , Deqing Jiang , Justin Sirignano

This paper introduces the Hamilton-Jacobi-Bellman Proximal Policy Optimization (HJBPPO) algorithm into reinforcement learning. The Hamilton-Jacobi-Bellman (HJB) equation is used in control theory to evaluate the optimality of the value…

Machine Learning · Computer Science 2023-02-02 Amartya Mukherjee , Jun Liu

Model-free reinforcement learning (RL) is a powerful approach for learning control policies directly from high-dimensional state and observation. However, it tends to be data-inefficient, which is especially costly in robotic learning…

Robotics · Computer Science 2020-10-14 Xubo Lyu , Mo Chen

This work addresses stochastic optimal control problems where the unknown state evolves in continuous time while partial, noisy, and possibly controllable measurements are only available in discrete time. We develop a framework for…

Optimization and Control · Mathematics 2025-08-19 Christian Bayer , Boualem Djehiche , Eliza Rezvanova , Raul Fidel Tempone