English
Related papers

Related papers: Policy Iteration Achieves Regularized Equilibrium …

200 papers

This paper presents a mathematical formulation to perform temporal parallelisation of continuous-time optimal control problems, which can be solved via the Hamilton--Jacobi--Bellman (HJB) equation. We divide the time interval of the control…

Optimization and Control · Mathematics 2024-12-18 Simo Särkkä , Ángel F. García-Fernández

A key problem in reinforcement learning for control with general function approximators (such as deep neural networks and other nonlinear functions) is that, for many algorithms employed in practice, updates to the policy or $Q$-function…

Machine Learning · Computer Science 2016-03-01 Joshua Achiam

We consider a stochastic optimal exit time feedback control problem. The Bellman equation is solved approximatively via the Policy Iteration algorithm on a polynomial ansatz space by a sequence of linear equations. As high degree…

Optimization and Control · Mathematics 2020-10-12 Konstantin Fackeldey , Mathias Oster , Leon Sallandt , Reinhold Schneider

We study the properties of the value function associated with an optimal control problem with uncertainties, known as average or Riemann-Stieltjes problem. Uncertainties are assumed to belong to a compact metric probability space, and…

Optimization and Control · Mathematics 2024-07-19 M. Soledad Aronna , Michele Palladino , Oscar Sierra

The partial stochastic realization of periodic processes from finite covariance data has recently been solved by Lindquist and Picci based on convex optimization of a generalized entropy functional. The meaning and the role of this…

Methodology · Statistics 2016-09-30 Giorgio Picci , Bin Zhu

This work presents a novel policy iteration algorithm to tackle nonzero-sum stochastic impulse games arising naturally in many applications. Despite the obvious impact of solving such problems, there are no suitable numerical methods…

Optimization and Control · Mathematics 2020-06-29 René Aïd , Francisco Bernal , Mohamed Mnif , Diego Zabaljauregui , Jorge P. Zubelli

We consider challenging dynamic programming models where the associated Bellman equation, and the value and policy iteration algorithms commonly exhibit complex and even pathological behavior. Our analysis is based on the new notion of…

Optimization and Control · Mathematics 2016-09-13 Dimitri P. Bertsekas

A class of linear kinetic Fokker-Planck equations with a non-trivial diffusion matrix and with periodic boundary conditions in the spatial variable is considered. After formulating the problem in a geometric setting, the question of the…

Mathematical Physics · Physics 2012-10-03 Simone Calogero

A classical approach for solving discrete time nonlinear control on a finite horizon consists in repeatedly minimizing linear quadratic approximations of the original problem around current candidate solutions. While widely popular in many…

Optimization and Control · Mathematics 2025-07-08 Vincent Roulet , Siddhartha Srinivasa , Maryam Fazel , Zaid Harchaoui

For pricing American options, %after suitable discretization in space and time, a sequence of discrete linear complementarity problems (LCPs) or equivalently Hamilton-Jacobi-Bellman (HJB) equations need to be solved in a sequential…

Numerical Analysis · Mathematics 2024-05-15 Xian-Ming Gu , Jun Liu , Cornelis W. Oosterlee

The standard version of the policy iteration (PI) algorithm fails for semicontinuous models, that is, for models with lower semicontinuous one-step costs and weakly continuous transition law. This is due to the lack of continuity properties…

Optimization and Control · Mathematics 2023-07-17 Óscar Vega-Amaya , Fernando Luque-Vásquez

We extend the construction of equilibria for linear-quadratic and mean-variance portfolio problems available in the literature to a large class of mean-field time-inconsistent stochastic control problems in continuous time. Our approach…

Optimization and Control · Mathematics 2021-10-01 Jiang Yu Nguwi , Nicolas Privault

The Ensemble Kalman inversion (EKI) method is a method for the estimation of unknown parameters in the context of (Bayesian) inverse problems. The method approximates the underlying measure by an ensemble of particles and iteratively…

Numerical Analysis · Mathematics 2021-08-02 Dirk Blömker , Claudia Schillings , Philipp Wacker , Simon Weissmann

We prove existence and uniqueness of stochastic equilibria in a class of incomplete continuous-time financial environments where the market participants are exponential utility maximizers with heterogeneous risk-aversion coefficients and…

General Finance · Quantitative Finance 2010-06-02 Gordan Zitkovic

Optimized certainty equivalents (OCEs) is a family of risk measures widely used by both practitioners and academics. This is mostly due to its tractability and the fact that it encompasses important examples, including entropic risk…

Optimization and Control · Mathematics 2022-06-07 Julio Backhoff Veraguas , A. Max Reppen , Ludovic Tangpi

We consider mean field social optimization in nonlinear diffusion models. By dynamic programming with a representative agent employing cooperative optimizer selection, we derive a new Hamilton--Jacobi--Bellman (HJB) equation to be called…

Optimization and Control · Mathematics 2026-05-19 Minyi Huang , Shuenn-Jyi Sheu , Li-Hsien Sun

An optimal control problem related to the probability of transition between stable states for a thermally driven Ginzburg-Landau equation is considered. The value function for the optimal control problem with a spatial discretization is…

Optimization and Control · Mathematics 2008-09-11 Mattias Sandberg

In optimal control problem, policy iteration (PI) is a powerful reinforcement learning (RL) tool used for designing optimal controller for the linear systems. However, the need for an initial stabilizing control policy significantly limits…

Optimization and Control · Mathematics 2024-11-13 Zhen Pang , Shengda Tang , Jun Cheng , Shuping He

The policy gradient theorem (Sutton et al., 2000) prescribes the usage of a cumulative discounted state distribution under the target policy to approximate the gradient. Most algorithms based on this theorem, in practice, break this…

Machine Learning · Computer Science 2022-07-08 Samuele Tosatto , Andrew Patterson , Martha White , A. Rupam Mahmood

We present a continuous-time equivalent to the well-known iterative linear-quadratic algorithm including an implementation of a backtracking line-search policy and a novel regularization approach based on the necessary conditions in the…

Systems and Control · Electrical Eng. & Systems 2025-05-22 Juraj Lieskovský , Jaroslav Bušek , Tomáš Vyhlídal