English
Related papers

Related papers: Policy Iteration for Exploratory Hamilton--Jacobi-…

200 papers

For pricing American options, %after suitable discretization in space and time, a sequence of discrete linear complementarity problems (LCPs) or equivalently Hamilton-Jacobi-Bellman (HJB) equations need to be solved in a sequential…

Numerical Analysis · Mathematics 2024-05-15 Xian-Ming Gu , Jun Liu , Cornelis W. Oosterlee

We consider a singular stochastic control problem, which is called the Monotone Follower Stochastic Control Problem and give sufficient conditions for the existence and uniqueness of a local-time type optimal control. To establish this…

Optimization and Control · Mathematics 2007-05-23 Erhan Bayraktar , Masahiko Egami

We develop a probabilistic framework for analysing model-based reinforcement learning in the episodic setting. We then apply it to study finite-time horizon stochastic control problems with linear dynamics but unknown coefficients and…

Machine Learning · Computer Science 2021-12-22 Lukasz Szpruch , Tanut Treetanthiploet , Yufei Zhang

We consider impulse control problems in finite horizon for diffusions with decision lag and execution delay. The new feature is that our general framework deals with the important case when several consecutive orders may be decided before…

Probability · Mathematics 2007-05-23 Benjamin Bruder , Huyen Pham

Equipping approximate dynamic programming (ADP) with inputconstraints has a tremendous significance. This enables ADP to be applied tothe systems with actuator limitations, which is quite common for dynamicalsystems. In a conventional…

Optimization and Control · Mathematics 2018-05-24 Xuefeng Bao , Zhi-Hong Mao , Nitin Sharma

We consider an optimal investment and consumption problem for a Black-Scholes financial market with stochastic coefficients driven by a diffusion process. We assume that an agent makes consumption and investment decisions based on CRRA…

Portfolio Management · Quantitative Finance 2011-12-12 Berdjane Belkacem , Serguei Pergamenchtchikov

We study a time-optimal control problem of a two-peakon collision. First, we state the controllability. Next, we find the time-optimal strategy. This is done via the HamiltonJacobi-Bellman equation and the dynamic programming method. We…

Optimization and Control · Mathematics 2021-05-24 Tomasz Cieślak , Bidesh Das

This paper characterizes differentiable and subgame Markov perfect equilibria in a continuous time intertemporal decision problem with non-constant discounting. Capturing the idea of non commitment by letting the commitment period being…

Optimization and Control · Mathematics 2008-08-29 Ivar Ekeland , Ali Lazrak

We study a stochastic optimal control problem for a partially observed diffusion. By using the control randomization method in [4], we prove a corresponding randomized dynamic programming principle (DPP) for the value function, which is…

Probability · Mathematics 2016-09-12 Elena Bandini , Andrea Cosso , Marco Fuhrman , Huyên Pham

Tackling large approximate dynamic programming or reinforcement learning problems requires methods that can exploit regularities, or intrinsic structure, of the problem in hand. Most current methods are geared towards exploiting the…

Machine Learning · Computer Science 2014-07-03 Amir-massoud Farahmand , Doina Precup , André M. S. Barreto , Mohammad Ghavamzadeh

We study continuous-time reinforcement learning (RL) for stochastic control in which system dynamics are governed by jump-diffusion processes. We formulate an entropy-regularized exploratory control problem with stochastic policies to…

Machine Learning · Computer Science 2025-08-26 Xuefeng Gao , Lingfei Li , Xun Yu Zhou

The goal of this article is to study fundamental mechanisms behind so-called indirect and direct data-driven control for unknown systems. Specifically, we consider policy iteration applied to the linear quadratic regulator problem. Two…

Systems and Control · Electrical Eng. & Systems 2024-04-30 Bowen Song , Andrea Iannelli

In this paper, we propose an interior-point method for linearly constrained optimization problems (possibly nonconvex). The method - which we call the Hessian barrier algorithm (HBA) - combines a forward Euler discretization of Hessian…

Optimization and Control · Mathematics 2023-09-14 Immanuel M. Bomze , Panayotis Mertikopoulos , Werner Schachinger , Mathias Staudigl

The numerical realization of the dynamic programming principle for continuous-time optimal control leads to nonlinear Hamilton-Jacobi-Bellman equations which require the minimization of a nonlinear mapping over the set of admissible…

Optimization and Control · Mathematics 2015-02-26 Dante Kalise , Axel Kröner , Karl Kunisch

We consider infinite-horizon $\gamma$-discounted Markov Decision Processes, for which it is known that there exists a stationary optimal policy. We consider the algorithm Value Iteration and the sequence of policies $\pi_1,...,\pi_k$ it…

Artificial Intelligence · Computer Science 2012-04-02 Bruno Scherrer

Under non-exponential discounting, we develop a dynamic theory for stopping problems in continuous time. Our framework covers discount functions that induce decreasing impatience. Due to the inherent time inconsistency, we look for…

Optimization and Control · Mathematics 2017-03-13 Yu-Jui Huang , Adrien Nguyen-Huu

We consider the infinite horizon risk-sensitive problem for nondegenerate diffusions with a compact action space, and controlled through the drift. We only impose a structural assumption on the running cost function, namely…

Optimization and Control · Mathematics 2019-03-20 Ari Arapostathis , Anup Biswas

We present discrete-time approximation of optimal control policies for infinite horizon discounted/ergodic control problems for controlled diffusions in $\Rd$\,. In particular, our objective is to show near optimality of optimal policies…

Optimization and Control · Mathematics 2025-02-11 Somnath Pradhan , Serdar Yuksel

We consider the homogenization of Hamilton-Jacobi equations and degenerate Bellman equations in stationary, ergodic, unbounded environments. We prove that, as the microscopic scale tends to zero, the equation averages to a deterministic…

Analysis of PDEs · Mathematics 2011-08-22 Scott N. Armstrong , Panagiotis E. Souganidis

We tackle the issue of finding a good policy when the number of policy updates is limited. This is done by approximating the expected policy reward as a sequence of concave lower bounds which can be efficiently maximized, drastically…

Artificial Intelligence · Computer Science 2016-12-30 Nicolas Le Roux
‹ Prev 1 8 9 10 Next ›