中文
相关论文

相关论文: Bridging Physics-Informed Neural Networks with Rei…

200 篇论文

We consider a deterministic optimal control problem with a maximum running cost functional, in a finite horizon context, and propose deep neural network approximations for Bellman's dynamic programming principle, corresponding also to some…

最优化与控制 · 数学 2022-10-11 Olivier Bokanowski , Xavier Warin , Averil Prost

This is the first in a series of papers in which we study an efficient approximation scheme for solving the Hamilton-Jacobi-Bellman equation for multi-dimensional problems in stochastic control theory. The method is a combination of a WKB…

计算金融 · 定量金融 2014-06-26 Sakda Chaiworawitkul , Patrick S. Hagan , Andrew Lesniewski

This paper considers consumption and portfolio optimization problems with recursive preferences in both infinite and finite time regions. Specially, the financial market consists of a risk-free asset and a risky asset that follows a general…

最优化与控制 · 数学 2024-12-30 Jian-hao Kang , Zhun Gou , Nan-jing Huang

Safe reinforcement learning aims to learn the optimal policy while satisfying safety constraints, which is essential in real-world applications. However, current algorithms still struggle for efficient policy updates with hard constraint…

机器学习 · 计算机科学 2022-06-20 Linrui Zhang , Li Shen , Long Yang , Shixiang Chen , Bo Yuan , Xueqian Wang , Dacheng Tao

The Hamilton Jacobi Bellman Equation (HJB) provides the globally optimal solution to large classes of control problems. Unfortunately, this generality comes at a price, the calculation of such solutions is typically intractible for systems…

最优化与控制 · 数学 2014-09-23 Matanya B. Horowitz , Anil Damle , Joel W. Burdick

This article approaches deterministic filtering via an application of the min-plus linearity of the corresponding dynamic programming operator. This filter design method yields a set-valued state estimator for discrete-time nonlinear…

最优化与控制 · 数学 2012-03-14 Abhijit G. Kallapur , Srinivas Sridharan , William M. McEneaney , Ian R. Petersen

Equipping approximate dynamic programming (ADP) with inputconstraints has a tremendous significance. This enables ADP to be applied tothe systems with actuator limitations, which is quite common for dynamicalsystems. In a conventional…

最优化与控制 · 数学 2018-05-24 Xuefeng Bao , Zhi-Hong Mao , Nitin Sharma

Classical reinforcement learning (RL) aims to optimize the expected cumulative rewards. In this work, we consider the RL setting where the goal is to optimize the quantile of the cumulative rewards. We parameterize the policy controlling…

机器学习 · 计算机科学 2022-02-17 Jinyang Jiang , Jiaqiao Hu , Yijie Peng

Neural network approaches that parameterize value functions have succeeded in approximating high-dimensional optimal feedback controllers when the Hamiltonian admits explicit formulas. However, many practical problems, such as the space…

最优化与控制 · 数学 2025-10-08 Eric Gelphman , Deepanshu Verma , Nicole Tianjiao Yang , Stanley Osher , Samy Wu Fung

In recent years, reinforcement learning (RL) has gained increasing attention in control engineering. Especially, policy gradient methods are widely used. In this work, we improve the tracking performance of proximal policy optimization…

We argue that inventory management presents unique opportunities for the reliable application of deep reinforcement learning (DRL). To enable this, we emphasize and test two complementary techniques. The first is Hindsight Differentiable…

机器学习 · 计算机科学 2025-09-12 Matias Alvo , Daniel Russo , Yash Kanoria , Minuk Lee

Quadratic Unconstrained Binary Optimization (QUBO) is a generic technique to model various NP-hard combinatorial optimization problems in the form of binary variables. The Hamiltonian function is often used to formulate QUBO problems where…

人工智能 · 计算机科学 2023-09-29 Redwan Ahmed Rizvee , Md. Mosaddek Khan

For continuous systems modeled by dynamical equations such as ODEs and SDEs, Bellman's Principle of Optimality takes the form of the Hamilton-Jacobi-Bellman (HJB) equation, which provides the theoretical target of reinforcement learning…

机器学习 · 计算机科学 2025-10-28 Haruki Settai , Naoya Takeishi , Takehisa Yairi

Designing optimal reward functions has been desired but extremely difficult in reinforcement learning (RL). When it comes to modern complex tasks, sophisticated reward functions are widely used to simplify policy learning yet even a tiny…

机器学习 · 计算机科学 2021-09-07 Ning Wei , Jiahua Liang , Di Xie , Shiliang Pu

When randomness in demand affects the sales of a product, retailers use dynamic pricing strategies to maximize their profits. In this article, we formulate the pricing problem as a continuous-time stochastic optimal control problem and find…

最优化与控制 · 数学 2019-03-13 Asbjørn Nilsen Riseth

Continuous-time stochastic processes underlie many natural and engineered systems. In healthcare, autonomous driving, and industrial control, direct interaction with the environment is often unsafe or impractical, motivating offline…

机器学习 · 统计学 2025-11-14 Nicolas Hoischen , Petar Bevanda , Max Beier , Stefan Sosnowski , Boris Houska , Sandra Hirche

This paper addresses distributional offline continuous-time reinforcement learning (DOCTR-L) with stochastic policies for high-dimensional optimal control. A soft distributional version of the classical Hamilton-Jacobi-Bellman (HJB)…

机器学习 · 计算机科学 2021-04-05 Igor Halperin

This paper proposes an off-policy risk-sensitive reinforcement learning based control framework for stabilization of a continuous-time nonlinear system that subjects to additive disturbances, input saturation, and state constraints. By…

系统与控制 · 电气工程与系统科学 2022-04-21 Cong Li , Qingchen Liu , Zhehua Zhou , Martin Buss , Fangzhou Liu

In this work, we provide theoretical guarantees for reward decomposition in deterministic MDPs. Reward decomposition is a special case of Hierarchical Reinforcement Learning, that allows one to learn many policies in parallel and combine…

机器学习 · 计算机科学 2018-03-14 Tom Zahavy , Avinatan Hasidim , Haim Kaplan , Yishay Mansour

Hard constraints in reinforcement learning (RL) often degrade policy performance. Lagrangian methods offer a way to blend objectives with constraints, but require intricate reward engineering and parameter tuning. In this work, we extend…

人工智能 · 计算机科学 2025-12-05 William Sharpless , Dylan Hirsch , Sander Tonkens , Nikhil Shinde , Sylvia Herbert