从生存核外学习:为何我们应构建能够优雅跌倒的机器人
人工智能
2018-06-19 v1 机器人学
摘要
尽管使用强化学习从零解决复杂问题取得了令人印象深刻的成果,但在机器人领域这仍主要局限于具有信息量很大奖励函数的基于模型的学习。主要挑战之一是奖励地形常存在大片无梯度的区域,使得难以有效采样梯度。我们在此表明,机器人状态初始化对奖励地形的影响可能比通常预期更为重要。特别地,我们展示了包含不可行初始化的反直觉益处,换言之,在以注定失败的状态中进行初始化。
引用
@article{arxiv.1806.06569,
title = {Learning from Outside the Viability Kernel: Why we Should Build Robots that can Fall with Grace},
author = {Steve Heim and Alexander Spröwitz},
journal= {arXiv preprint arXiv:1806.06569},
year = {2018}
}
备注
Proceedings of the 2018 IEEE International Conference on SImulation, Modeling and Programming for Autonomous Robots (SIMPAR), Brisbane, Australia, 16-19 2018