English

USHER: Unbiased Sampling for Hindsight Experience Replay

Machine Learning 2022-07-05 v1 Artificial Intelligence Robotics

Abstract

Dealing with sparse rewards is a long-standing challenge in reinforcement learning (RL). Hindsight Experience Replay (HER) addresses this problem by reusing failed trajectories for one goal as successful trajectories for another. This allows for both a minimum density of reward and for generalization across multiple goals. However, this strategy is known to result in a biased value function, as the update rule underestimates the likelihood of bad outcomes in a stochastic environment. We propose an asymptotically unbiased importance-sampling-based algorithm to address this problem without sacrificing performance on deterministic environments. We show its effectiveness on a range of robotic systems, including challenging high dimensional stochastic environments.

Keywords

Cite

@article{arxiv.2207.01115,
  title  = {USHER: Unbiased Sampling for Hindsight Experience Replay},
  author = {Liam Schramm and Yunfu Deng and Edgar Granados and Abdeslam Boularias},
  journal= {arXiv preprint arXiv:2207.01115},
  year   = {2022}
}
R2 v1 2026-06-24T12:12:36.340Z