An Empirical Dynamic Programming Algorithm for Continuous MDPs

William B. Haskell; Rahul Jain; Hiteshi Sharma; Pengqian Yu

An Empirical Dynamic Programming Algorithm for Continuous MDPs

Optimization and Control 2019-04-25 v2

Authors: William B. Haskell , Rahul Jain , Hiteshi Sharma , Pengqian Yu

Abstract

We propose universal randomized function approximation-based empirical value iteration (EVI) algorithms for Markov decision processes. The `empirical' nature comes from each iteration being done empirically from samples available from simulations of the next state. This makes the Bellman operator a random operator. A parametric and a non-parametric method for function approximation using a parametric function space and the Reproducing Kernel Hilbert Space (RKHS) respectively are then combined with EVI. Both function spaces have the universal function approximation property. Basis functions are picked randomly. Convergence analysis is done using a random operator framework with techniques from the theory of stochastic dominance. Finite time sample complexity bounds are derived for both universal approximate dynamic programming algorithms. Numerical experiments support the versatility and effectiveness of this approach.

Keywords

markov decision processes markov chain monte carlo dynamic programming

Cite

@article{arxiv.1709.07506,
  title  = {An Empirical Dynamic Programming Algorithm for Continuous MDPs},
  author = {William B. Haskell and Rahul Jain and Hiteshi Sharma and Pengqian Yu},
  journal= {arXiv preprint arXiv:1709.07506},
  year   = {2019}
}

Comments

Accepted for publication in IEEE Transactions on Automatic Control

An Empirical Dynamic Programming Algorithm for Continuous MDPs

Abstract

Keywords

Cite

Comments

Related papers