无需在线探索学习鲁棒驾驶策略
机器人学
2021-03-16 v1
摘要
我们提出一种多时间尺度预测表示学习方法,以离线方式高效学习鲁棒驾驶策略,使其能很好地泛化到离线训练数据中未涵盖的新颖道路几何形状,以及受损和干扰性的车道条件。我们展示了所提出的表示学习方法可轻松应用于离线(批量)强化学习设定,与标准批量RL方法相比,在新颖条件下表现出良好的泛化能力与效率。我们提出的方法利用完全在真实世界中离线收集的训练数据,从而消除了密集在线探索的需求,后者阻碍了深度强化学习在真实世界机器人训练中的应用。我们在模拟器和真实世界场景中进行了多种实验,以评估和分析我们所提出的论断。
引用
@article{arxiv.2103.08070,
title = {Learning robust driving policies without online exploration},
author = {Daniel Graves and Nhat M. Nguyen and Kimia Hassanzadeh and Jun Jin and Jun Luo},
journal= {arXiv preprint arXiv:2103.08070},
year = {2021}
}
备注
Accepted in ICRA 2021. Due to format limitations of ICRA, we include appendix of our detailed evaluation results in this full version. arXiv admin note: substantial text overlap with arXiv:2006.15110