English

Mixing Human Demonstrations with Self-Exploration in Experience Replay for Deep Reinforcement Learning

Artificial Intelligence 2021-07-15 v1

Abstract

We investigate the effect of using human demonstration data in the replay buffer for Deep Reinforcement Learning. We use a policy gradient method with a modified experience replay buffer where a human demonstration experience is sampled with a given probability. We analyze different ratios of using demonstration data in a task where an agent attempts to reach a goal while avoiding obstacles. Our results suggest that while the agents trained by pure self-exploration and pure demonstration had similar success rates, the pure demonstration model converged faster to solutions with less number of steps.

Keywords

Cite

@article{arxiv.2107.06840,
  title  = {Mixing Human Demonstrations with Self-Exploration in Experience Replay for Deep Reinforcement Learning},
  author = {Dylan Klein and Akansel Cosgun},
  journal= {arXiv preprint arXiv:2107.06840},
  year   = {2021}
}

Comments

2 pages. Submitted to ICDL 2021 Workshop on Human aligned Reinforcement Learning for Autonomous Agents and Robots