English

Approximation Benefits of Policy Gradient Methods with Aggregated States

Machine Learning 2022-06-24 v3 Optimization and Control Machine Learning

Abstract

Folklore suggests that policy gradient can be more robust to misspecification than its relative, approximate policy iteration. This paper studies the case of state-aggregated representations, where the state space is partitioned and either the policy or value function approximation is held constant over partitions. This paper shows a policy gradient method converges to a policy whose regret per-period is bounded by ϵ\epsilon, the largest difference between two elements of the state-action value function belonging to a common partition. With the same representation, both approximate policy iteration and approximate value iteration can produce policies whose per-period regret scales as ϵ/(1γ)\epsilon/(1-\gamma), where γ\gamma is a discount factor. Faced with inherent approximation error, methods that locally optimize the true decision-objective can be far more robust.

Keywords

Cite

@article{arxiv.2007.11684,
  title  = {Approximation Benefits of Policy Gradient Methods with Aggregated States},
  author = {Daniel Russo},
  journal= {arXiv preprint arXiv:2007.11684},
  year   = {2022}
}
R2 v1 2026-06-23T17:19:47.899Z