English

Policy Iteration for Exploratory Hamilton--Jacobi--Bellman Equations

Optimization and Control 2025-05-28 v4 Analysis of PDEs

Abstract

We study the policy iteration algorithm (PIA) for entropy-regularized stochastic control problems on an infinite time horizon with a large discount rate, focusing on two main scenarios. First, we analyze PIA with bounded coefficients where the controls applied to the diffusion term satisfy a smallness condition. We demonstrate the convergence of PIA based on a uniform C2,α\mathcal{C}^{2,\alpha} estimate for the value sequence generated by PIA, and provide a quantitative convergence analysis for this scenario. Second, we investigate PIA with unbounded coefficients but no control over the diffusion term. In this scenario, we first provide the well-posedness of the exploratory Hamilton--Jacobi--Bellman equation with linear growth coefficients and polynomial growth reward function. By such a well-posedess result we achieve PIA's convergence by establishing a quantitative locally uniform C1,α\mathcal{C}^{1,\alpha} estimates for the generated value sequence.

Keywords

Cite

@article{arxiv.2406.00612,
  title  = {Policy Iteration for Exploratory Hamilton--Jacobi--Bellman Equations},
  author = {Hung Vinh Tran and Zhenhua Wang and Yuming Paul Zhang},
  journal= {arXiv preprint arXiv:2406.00612},
  year   = {2025}
}

Comments

25 pages

R2 v1 2026-06-28T16:49:52.678Z