中文
相关论文

相关论文: Continuous-Time Fitted Value Iteration for Robust …

200 篇论文

We consider the portfolio optimisation problem where the terminal function is an S-shaped utility applied at the difference between the wealth and a random benchmark process. We develop several numerical methods for solving the problem…

计算金融 · 定量金融 2024-10-10 Ashley Davey , Harry Zheng

Safe reinforcement learning (RL) aims to learn policies that satisfy certain constraints before deploying them to safety-critical applications. Previous primal-dual style approaches suffer from instability issues and lack optimality…

机器学习 · 计算机科学 2022-06-20 Zuxin Liu , Zhepeng Cen , Vladislav Isenbaev , Wei Liu , Zhiwei Steven Wu , Bo Li , Ding Zhao

In this paper, we study a stochastic recursive optimal control problem in which the system is governed by a functional forward-backward stochastic differential equation. Under standard assumptions, we establish the dynamic programming…

概率论 · 数学 2013-01-03 Shaolin Ji , Shuzhen Yang

The control of relaxation-type systems of ordinary differential equations is investigated using the Hamilton-Jacobi-Bellman equation. First, we recast the model as a singularly perturbed dynamics which we embed in a family of controlled…

最优化与控制 · 数学 2024-04-23 Michael Herty , Hicham Kouhkouh

Reinforcement learning algorithms have shown great success in solving different problems ranging from playing video games to robotics. However, they struggle to solve delicate robotic problems, especially those involving contact…

机器人学 · 计算机科学 2020-07-15 Miroslav Bogdanovic , Majid Khadiv , Ludovic Righetti

This paper presents a novel model-free Reinforcement Learning algorithm for learning behavior in continuous action, state, and goal spaces. The algorithm approximates optimal value functions using non-parametric estimators. It is able to…

人工智能 · 计算机科学 2019-08-28 Andreas Gerken , Michael Spranger

We study a class of optimal control problems with state constraints where the state equation is a differential equation with delays. This class includes some problems arising in economics, in particular the so-called models with time to…

最优化与控制 · 数学 2009-07-09 Salvatore Federico , Ben Goldys , Fausto Gozzi

The goal of this thesis is to provide efficient and provably convergent numerical methods for solving partial differential equations (PDEs) coming from impulse control problems motivated by finance. Impulses, which are controlled jumps in a…

数值分析 · 数学 2018-02-05 Parsiad Azimzadeh

The paper introduces an interactive machine learning mechanism to process the measurements of an uncertain, nonlinear dynamic process and hence advise an actuation strategy in real-time. For concept demonstration, a trajectory-following…

系统与控制 · 电气工程与系统科学 2023-03-16 Mohammed Abouheaf , Derek Boase , Wail Gueaieb , Davide Spinello , Salah Al-Sharhan

Keeping risk under control is often more crucial than maximizing expected rewards in real-world decision-making situations, such as finance, robotics, autonomous driving, etc. The most natural choice of risk measures is variance, which…

机器学习 · 计算机科学 2023-03-09 Xiaoteng Ma , Shuai Ma , Li Xia , Qianchuan Zhao

An off policy reinforcement learning based control strategy is developed for the optimal tracking control problem to achieve the prescribed performance of full states during the learning process. The optimal tracking control problem is…

系统与控制 · 电气工程与系统科学 2020-09-02 C. Li , Y. Wang , F. Liu , M. Buss

We study an inverse problem of the stochastic optimal control of general diffusions with performance index having the quadratic penalty term of the control process. Under mild conditions on the system dynamics, the cost functions, and the…

最优化与控制 · 数学 2022-11-17 Yumiharu Nakano

In this manuscript, we study optimal control problems for stochastic delay differential equations using the dynamic programming approach in Hilbert spaces via viscosity solutions of the associated Hamilton-Jacobi-Bellman equations. We show…

最优化与控制 · 数学 2024-12-24 Filippo de Feo , Andrzej Święch

In this paper, we study one kind of stochastic recursive optimal control problem with the obstacle constraints for the cost function where the cost function is described by the solution of one reflected backward stochastic differential…

最优化与控制 · 数学 2007-05-23 Zhen Wu , Zhiyong Yu

This paper considers consumption and portfolio optimization problems with recursive preferences in both infinite and finite time regions. Specially, the financial market consists of a risk-free asset and a risky asset that follows a general…

最优化与控制 · 数学 2024-12-30 Jian-hao Kang , Zhun Gou , Nan-jing Huang

We study sequential decision making in environments where rewards are only partially observed, but can be modeled as a function of observed contexts and the chosen action by the decision maker. This setting, known as contextual bandits,…

统计方法学 · 统计学 2015-03-11 Miroslav Dudík , Dumitru Erhan , John Langford , Lihong Li

Many optimal control problems are formulated as two point boundary value problems (TPBVPs) with conditions of optimality derived from the Hamilton-Jacobi-Bellman (HJB) equations. In most cases, it is challenging to solve HJBs due to the…

最优化与控制 · 数学 2019-07-25 Sixiong You , Ran Dai , Ping Lu

Reinforcement Learning is a powerful framework for training agents to navigate different situations, but it is susceptible to changes in environmental dynamics. However, solving Markov Decision Processes that are robust to changes is…

机器学习 · 计算机科学 2024-06-21 Etash Kumar Guha

Real-world applications require RL algorithms to act safely. During learning process, it is likely that the agent executes sub-optimal actions that may lead to unsafe/poor states of the system. Exploration is particularly brittle in…

机器学习 · 统计学 2019-06-17 Elena Smirnova , Elvis Dohmatob , Jérémie Mary

This paper studies an optimal stochastic impulse control problem in a finite horizon with a decision lag, by which we mean that after an impulse is made, a fixed number units of time has to be elapsed before the next impulse is allowed to…

最优化与控制 · 数学 2021-02-09 Chang Li , Jiongmin Yong