中文

通过玩耍学习——从零开始解决稀疏奖励任务

机器学习 2018-03-01 v1 机器人学 机器学习

摘要

我们提出调度辅助控制(Scheduled Auxiliary Control, SAC-X),一种在强化学习(RL)背景下的新学习范式。SAC-X能够在存在多个稀疏奖励信号的情况下——从零开始——学习复杂行为。为此,智能体配备了一组通用辅助任务,它通过离策略RL尝试同时学习这些任务。我们方法背后的关键思想是,对辅助策略的主动(习得)调度与执行使智能体能够高效探索其环境——使其擅长稀疏奖励RL。我们在若干具有挑战性的机器人操纵设定中的实验展示了我们方法的威力。

关键词

引用

@article{arxiv.1802.10567,
  title  = {Learning by Playing - Solving Sparse Reward Tasks from Scratch},
  author = {Martin Riedmiller and Roland Hafner and Thomas Lampe and Michael Neunert and Jonas Degrave and Tom Van de Wiele and Volodymyr Mnih and Nicolas Heess and Jost Tobias Springenberg},
  journal= {arXiv preprint arXiv:1802.10567},
  year   = {2018}
}

备注

A video of the rich set of learned behaviours can be found at https://youtu.be/mPKyvocNe_M