反应式行为反馈模型的监督学习与强化学习:触觉反馈试验台
机器人学
2022-12-06 v2
摘要
机器人需要能够适应环境中意想不到的变化,从而自主地完成其任务。然而,手工设计用于适应的反馈模型是繁琐的(即便可能),这使得数据驱动方法成为一种有前景的替代方案。本文介绍了一个用于学习反应式运动规划反馈模型的完整框架。我们的流程首先通过半自动分割算法将完整任务的演示分割为运动基元。然后,给定额外的成功适应行为演示,我们通过示范学习(learning from demonstrations)来学习初始反馈模型。在最后阶段,一种样本高效的强化学习算法通过少量真实系统交互,针对新颖任务设置对这些反馈模型进行微调。我们在真实拟人机器人上通过学习触觉反馈任务来评估我们的方法。
引用
@article{arxiv.2007.00450,
title = {Supervised Learning and Reinforcement Learning of Feedback Models for Reactive Behaviors: Tactile Feedback Testbed},
author = {Giovanni Sutanto and Katharina Rombach and Yevgen Chebotar and Zhe Su and Stefan Schaal and Gaurav S. Sukhatme and Franziska Meier},
journal= {arXiv preprint arXiv:2007.00450},
year = {2022}
}
备注
Accepted for publication in the International Journal of Robotics Research (IJRR). Paper length is 22 pages (including references) with 12 figures. A video overview of the reinforcement learning experiment on the real robot can be seen at https://www.youtube.com/watch?v=yu5v-ZXo4-E. arXiv admin note: text overlap with arXiv:1710.08555