面向机器人音乐家的带时间依赖目标强化学习
机器人学
2020-11-12 v1 人工智能
摘要
强化学习是一种很有前景的机器人控制任务实现方法。然而,演奏乐器这一任务在很大程度上未被探索,因为它涉及实现具有时间维度的序列目标——旋律——这一挑战。在本文中,我们通过引入目标条件强化学习的时间扩展:时间依赖目标,来解决机器人音乐演奏问题。我们证明了这些可用于训练机器人音乐家演奏特雷门琴。我们在仿真中训练机器人智能体,并将习得的策略迁移到真实世界的机器人特雷门琴演奏者上。补充视频:https://youtu.be/jvC9mPzdQN4
引用
@article{arxiv.2011.05715,
title = {Reinforcement Learning with Time-dependent Goals for Robotic Musicians},
author = {Thilo Fryen and Manfred Eppe and Phuong D. H. Nguyen and Timo Gerkmann and Stefan Wermter},
journal= {arXiv preprint arXiv:2011.05715},
year = {2020}
}
备注
Preprint, submitted to IEEE Robotics and Automation Letters (RA-L) 2021 with International Conference on Robotics and Automation Conference Option (ICRA) 2021