English
Related papers

Related papers: Policy Iteration Achieves Regularized Equilibrium …

200 papers

This paper concerns continuous dependence estimates for Hamilton-Jacobi-Bellman-Isaacs operators (briefly, HJBI). For the parabolic Cauchy problem, we establish such an estimate in the whole space $[0,+\infty)\times\Rn$. Moreover, under…

Analysis of PDEs · Mathematics 2010-08-02 Claudio Marchi

This paper deals with the long time behavior of solutions to the spatially homogeneous Boltzmann equation. The interactions considered are the so-called (non cut-off and non mollified) hard potentials. We prove an exponential in time…

Analysis of PDEs · Mathematics 2015-12-22 Isabelle Tristani

This is a companion paper of [Mixed equilibrium solution of time-inconsistent stochastic LQ problem, arXiv:1802.03032], where general theory has been established to characterize the open-loop equilibrium control, feedback equilibrium…

Optimization and Control · Mathematics 2018-03-26 Yuan-Hua Ni , Xun Li , Ji-Feng Zhang , Miroslav Krstic

This paper considers an infinite-horizon Markov decision process (MDP) that allows for general non-exponential discount functions, in both discrete and continuous time. Due to the inherent time inconsistency, we look for a randomized…

Optimization and Control · Mathematics 2024-12-10 Erhan Bayraktar , Yu-Jui Huang , Zhenhua Wang , Zhou Zhou

The Hamilton-Jacobi-Bellman equation (HJB) associated with the time inhomogeneous singular control problem is a parabolic partial differential equation, and the existence of a classical solution is usually difficult to prove. In this paper,…

Optimization and Control · Mathematics 2014-10-14 Yipeng Yang

We introduce a contractive abstract dynamic programming framework and related policy iteration algorithms, specifically designed for sequential zero-sum games and minimax problems with a general structure. Aside from greater generality, the…

Computer Science and Game Theory · Computer Science 2021-10-22 Dimitri Bertsekas

In this paper, we focus on a class of time-inconsistent stochastic control problems, where the objective function includes the mean and several higher-order central moments of the terminal value of state. To tackle the time-inconsistency,…

Mathematical Finance · Quantitative Finance 2025-05-08 Yike Wang , Jingzhen Liu , Alain Bensoussan , Ka-Fai Cedric Yiu , Jiaqin Wei

Tackling large approximate dynamic programming or reinforcement learning problems requires methods that can exploit regularities, or intrinsic structure, of the problem in hand. Most current methods are geared towards exploiting the…

Machine Learning · Computer Science 2014-07-03 Amir-massoud Farahmand , Doina Precup , André M. S. Barreto , Mohammad Ghavamzadeh

This paper investigates a continuous-time portfolio optimization problem with the following features: (i) a no-short selling constraint; (ii) a leverage constraint, that is, an upper limit for the sum of portfolio weights; and (iii) a…

Portfolio Management · Quantitative Finance 2022-03-08 Masashi Ieda

An exponentially convergent numerical method for solving a differential equation with a right-hand fractional Riemann-Liouville time-derivative and an unbounded operator coefficient in Banach space is proposed and analysed for a…

Numerical Analysis · Mathematics 2024-12-24 V. Vasylyk , V. L. Makarov

We study the problem of computing the value function from a discretely-observed trajectory of a continuous-time diffusion process. We develop a new class of algorithms based on easily implementable numerical schemes that are compatible with…

Machine Learning · Computer Science 2024-07-09 Wenlong Mou , Yuhua Zhu

The goal of this paper is to investigate new and simple convergence analysis of dynamic programming for linear quadratic regulator problem of discrete-time linear time-invariant systems. In particular, bounds on errors are given in terms of…

Optimization and Control · Mathematics 2021-06-18 Donghwan Lee

In this paper, we focus on the stochastic representation of a system of coupled Hamilton-Jacobi-Bellman-Isaacs (HJB-Isaacs (HJBI), for short) equations which is in fact a system of coupled Isaacs' type integral-partial differential…

Optimization and Control · Mathematics 2023-07-12 Sheng Luo , Wenqiang Li , Xun Li , Qingmeng Wei

In this paper we propose the design of an iterative observer using space as a time-like variable and prove its convergence. The iterative observer algorithm solves boundary estimation problem for a steady-state elliptic equation system…

Numerical Analysis · Mathematics 2016-04-22 Muhammad Usman Majeed , Taous Meriem Laleg-Kirati

Modified policy iteration (MPI) is a dynamic programming algorithm that combines elements of policy iteration and value iteration. The convergence of MPI has been well studied in the context of discounted and average-cost MDPs. In this…

Machine Learning · Computer Science 2024-02-16 Yashaswini Murthy , Mehrdad Moharrami , R. Srikant

Decision-making problems in uncertain or stochastic domains are often formulated as Markov decision processes (MDPs). Policy iteration (PI) is a popular algorithm for searching over policy-space, the size of which is exponential in the…

Artificial Intelligence · Computer Science 2013-01-30 Yishay Mansour , Satinder Singh

We consider the infinite-horizon discounted optimal control problem formalized by Markov Decision Processes. We focus on several approximate variations of the Policy Iteration algorithm: Approximate Policy Iteration, Conservative Policy…

Artificial Intelligence · Computer Science 2014-05-13 Bruno Scherrer

For continuous systems modeled by dynamical equations such as ODEs and SDEs, Bellman's Principle of Optimality takes the form of the Hamilton-Jacobi-Bellman (HJB) equation, which provides the theoretical target of reinforcement learning…

Machine Learning · Computer Science 2025-10-28 Haruki Settai , Naoya Takeishi , Takehisa Yairi

In this paper, we present a scalable deep learning approach to solve opinion dynamics stochastic optimal control problems with mean field term coupling in the dynamics and cost function. Our approach relies on the probabilistic…

Multiagent Systems · Computer Science 2022-04-19 Tianrong Chen , Ziyi Wang , Evangelos A. Theodorou

The semi-implicit Euler-Maruyama (EM) method is investigated to approximate a class of time-changed stochastic differential equations, whose drift coefficient can grow super-linearly and diffusion coefficient obeys the global Lipschitz…

Numerical Analysis · Mathematics 2019-07-29 Chang-Song Deng , Wei Liu