中文

RobotKeyframing:通过稠密稀疏奖励混合学习 locomotion 中的高级目标关键帧

机器人学 2024-11-05 v2 人工智能 机器学习

摘要

本文提出了一种新颖的基于学习的控制框架,使用关键帧方法在四足机器人的自然 locomotion 中融合高级目标。这些高级目标指定为可变数量的部分或完整姿态目标,以任意时间间隔分布。我们提出的框架利用多 critic 强化学习算法有效处理稠密稀疏奖励的混合情况。此外,它采用基于 transformer 的编码器以 accommodate 可变数量的输入目标,每个目标都与特定的时间到达关联。在仿真和硬件实验期间,我们证明了该框架能够在所需时间满足目标关键帧序列。实验中,多 critic 方法显著减少了相对于标准单 critic 方法的超参数调优工作量。此外,提出的基于 transformer 的体系结构使机器人能够预见未来目标,从而在其到达目标能力方面实现定量改进。

关键词

引用

@article{arxiv.2407.11562,
  title  = {RobotKeyframing: Learning Locomotion with High-Level Objectives via Mixture of Dense and Sparse Rewards},
  author = {Fatemeh Zargarbashi and Jin Cheng and Dongho Kang and Robert Sumner and Stelian Coros},
  journal= {arXiv preprint arXiv:2407.11562},
  year   = {2024}
}

备注

This paper has been accepted to 8th Conference on Robot Learning (CoRL 2024). Project website: https://sites.google.com/view/robot-keyframing