中文

模仿学习的超参数选择

机器学习 2021-05-26 v1

摘要

我们解决在连续控制背景下模仿学习算法的超参数(HPs)调优问题,其中示范专家的基础奖励函数在任何时刻均不可观测。模仿学习的大量文献大多假定该奖励函数可用于HP选择,但这并非现实设定。事实上,若该奖励函数可用,便可直接用于策略训练,模仿也就不必要了。为应对这一被大多忽略的问题,我们提出若干外部奖励的可能代理。我们在一个广泛的实证研究中(跨越9个环境、超过10'000个智能体)对其评估,并给出选择HPs的实用建议。我们的结果表明,尽管模仿学习算法对HP选择敏感,但通常可通过奖励函数的代理选出足够好的HPs。

关键词

引用

@article{arxiv.2105.12034,
  title  = {Hyperparameter Selection for Imitation Learning},
  author = {Leonard Hussenot and Marcin Andrychowicz and Damien Vincent and Robert Dadashi and Anton Raichuk and Lukasz Stafiniak and Sertan Girgin and Raphael Marinier and Nikola Momchev and Sabela Ramos and Manu Orsini and Olivier Bachem and Matthieu Geist and Olivier Pietquin},
  journal= {arXiv preprint arXiv:2105.12034},
  year   = {2021}
}

备注

ICML 2021