English
Related papers

Related papers: A Temporal Difference Method for Stochastic Contin…

200 papers

We present a novel method to approximate optimal feedback laws for nonlinear optimal control based on low-rank tensor train (TT) decompositions. The approach is based on the Dirac-Frenkel variational principle with the modification that the…

Optimization and Control · Mathematics 2021-11-30 Martin Eigel , Reinhold Schneider , David Sommer

Environmental management optimizing a long-run objective is an ergodic control problem whose resolution can be achieved by solving an associated non-local Hamilton-Jacobi-Bellman (HJB) equation having an effective Hamiltonian. Focusing on…

Optimization and Control · Mathematics 2022-05-11 Hidekazu Yoshioka , Motoh Tsujimura , Yuta Yaegashi

In this article, a class of optimal control problems of differential equations with delays are investigated for which the associated Hamilton-Jacobi-Bellman (HJB) equations are nonlinear partial differential equations with delays. This type…

Optimization and Control · Mathematics 2015-07-16 Jianjun Zhou

We consider the problem of learning a set of probability distributions from the empirical Bellman dynamics in distributional reinforcement learning (RL), a class of state-of-the-art methods that estimate the distribution, as opposed to only…

Machine Learning · Computer Science 2020-12-10 Thanh Tang Nguyen , Sunil Gupta , Svetha Venkatesh

We study a stochastic control problem on a bounded domain, which arises from a continuous-time optimal management model. Via the corresponding Hamilton-Jacobi-Bellman equation the value function is shown to be jointly continuous and to…

Probability · Mathematics 2017-10-24 Ruoting Gong , Christian Houdré

We propose a model-based reinforcement learning (RL) approach for noisy time-dependent gate optimization with improved sample complexity over model-free RL. Sample complexity is the number of controller interactions with the physical…

This paper studies an optimal stochastic impulse control problem in a finite horizon with a decision lag, by which we mean that after an impulse is made, a fixed number units of time has to be elapsed before the next impulse is allowed to…

Optimization and Control · Mathematics 2021-02-09 Chang Li , Jiongmin Yong

Reward specification plays a central role in reinforcement learning (RL), guiding the agent's behavior. To express non-Markovian rewards, formalisms such as reward machines have been introduced to capture dependencies on histories. However,…

Artificial Intelligence · Computer Science 2026-05-13 Rajarshi Roy , Anirban Majumdar , Ritam Raha , David Parker , Marta Kwiatkowska

In this paper we study a first extension of the theory of mild solutions for HJB equations in Hilbert spaces to the case when the domain is not the whole space. More precisely, we consider a half-space as domain, and a semilinear…

Optimization and Control · Mathematics 2022-09-30 Alessandro Calvia , Gianluca Cappa , Fausto Gozzi , Enrico Priola

The convergence of many reinforcement learning (RL) algorithms with linear function approximation has been investigated extensively but most proofs assume that these methods converge to a unique solution. In this paper, we provide a…

Machine Learning · Computer Science 2019-05-29 Marcus Hutter , Samuel Yang-Zhao , Sultan J. Majeed

Solving the Hamilton-Jacobi-Bellman equation is important in many domains including control, robotics and economics. Especially for continuous control, solving this differential equation and its extension the Hamilton-Jacobi-Isaacs…

Robotics · Computer Science 2021-10-06 Michael Lutter , Boris Belousov , Shie Mannor , Dieter Fox , Animesh Garg , Jan Peters

We propose a new numerical method for solving the Hamilton-Jacobi-Bellman quasi-variational inequality associated with the combined impulse and stochastic optimal control problem over a finite time horizon. Our method corresponds to an…

Numerical Analysis · Mathematics 2015-02-05 Masashi Ieda

We investigate an optimal control problem for a diffusion whose drift and running cost are merely measurable in the state variable. Such low regularity rules out the use of Pontryagin's maximum principle and also invalidates the standard…

Optimization and Control · Mathematics 2025-09-03 Kai Du , Qingmeng Wei

We study reinforcement learning (RL) in the setting of continuous time and space, for an infinite horizon with a discounted objective and the underlying dynamics driven by a stochastic differential equation. Built upon recent advances in…

Machine Learning · Computer Science 2023-10-19 Hanyang Zhao , Wenpin Tang , David D. Yao

In this article, we provide a numerical method based on fitted finite volume method to approximate the Hamilton-Jacobi-Bellman (HJB) equation coming from stochastic optimal control problems. The computational challenge is due to the nature…

Numerical Analysis · Mathematics 2020-02-21 Christelle Dleuna Nyoumbi , Antoine Tambue

The objective of designing a control system is to steer a dynamical system with a control signal, guiding it to exhibit the desired behavior. The Hamilton-Jacobi-Bellman (HJB) partial differential equation offers a framework for optimal…

Machine Learning · Computer Science 2025-10-22 Jostein Barry-Straume , Adwait D. Verulkar , Arash Sarshar , Andrey A. Popov , Adrian Sandu

In order to obviate the requirement of drift dynamics in adaptive dynamic programming (ADP), integral reinforcement learning (IRL) has been proposed as an alternate formulation of Bellman equation.However control coupling dynamics is still…

Systems and Control · Electrical Eng. & Systems 2020-05-11 Amardeep Mishra , Satadal Ghosh

This paper presents a learning-based optimal control framework for safety-critical systems with parametric uncertainties, addressing both time-triggered and self-triggered controller implementations. First, we develop a robust control…

Systems and Control · Electrical Eng. & Systems 2025-07-31 Zhanglin Shangguan , Bo Yang , Qi Li , Wei Xiao , Xingping Guan

We consider a Bolza-type optimal control problem for a dynamical system described by a fractional differential equation with the Caputo derivative of an order $\alpha \in (0, 1)$. The value of this problem is introduced as a functional in a…

Optimization and Control · Mathematics 2019-08-06 Mikhail I. Gomoyunov

Model-reference adaptive systems refer to a consortium of techniques that guide plants to track desired reference trajectories. Approaches based on theories like Lyapunov, sliding surfaces, and backstepping are typically employed to advise…

Systems and Control · Electrical Eng. & Systems 2023-03-20 Mohammed Abouheaf , Wail Gueaieb , Davide Spinello , Salah Al-Sharhan