探索性模仿学习:基于路径符号的连续环境方法
机器学习
2024-07-23 v2 人工智能
摘要
一些模仿学习方法将行为克隆与自我监督结合,从状态对中推断动作。然而,这些方法大多依赖大量专家轨迹以提高泛化性,以及人类干预以捕捉问题的关键方面,如域约束。在本文中,我们提出了 Continuous Imitation Learning from Observation(CILO),一种新的模仿学习方法,具备两个重要特征:(i)探索,允许更多样的状态转换,需 fewer(需求更少)专家轨迹,导致更少的训练迭代;(ii)路径符号(path signatures),通过创建代理和专家轨迹的非参数表示,实现自动约束编码。我们将 CILO 与基线方法以及两种领先的模仿学习方法进行了比较,测试于五个环境中。在所有环境中,CILO 的整体性能均优于其他方法,在两个环境中还超过了专家。
引用
@article{arxiv.2407.04856,
title = {Explorative Imitation Learning: A Path Signature Approach for Continuous Environments},
author = {Nathan Gavenski and Juarez Monteiro and Felipe Meneguzzi and Michael Luck and Odinaldo Rodrigues},
journal= {arXiv preprint arXiv:2407.04856},
year = {2024}
}
备注
This paper has been accepted in the 27th European Conference on Artificial Intelligence (ECAI) 2024