CUPID:基于影响函数 curate robot data
机器人学
2025-09-25 v2 人工智能
机器学习
摘要
在机器人模仿学习中,策略性能严格取决于演示数据的数据质量和组成。然而,开发对 individual 演示如何贡献于下游结果(如闭环任务成功或失败)的精确理解,仍是一个持久的挑战。我们提出 CUPID,一种基于 novel 情况函数理论化� formulation for imitation learning policies 的机器人数据 curate 方法。给定一组评估 rollout,CUPID 估计每个训练演示对 policy 预期回报的影响。这 enables 根据 policy 闭环性能的影响对演示进行排序和选择。我们利用 CUPID 通过 1) 滤除损害 policy 性能的训练演示,以及 2) 从新收集的轨迹中挑选最能提升 policy 的轨迹来进行数据 curate。大量仿真和硬件实验表明,我们的方法始终能够识别哪些数据驱动测试时性能。例如,使用经 CUPID curate 的数据量不足原数据的 33% 即可在仿真 RoboMimic benchmark 上实现 state-of-the-art diffusion policies,硬件实验中也取得了类似的收益。此外,硬件实验表明,我们的方法能够识别在分布迁移下稳健的策略、隔离虚假关联,甚至提升通用机器人 policy 的 post-training 效果。视频和代码已公开于:https://cupid-curation.github.io。
引用
@article{arxiv.2506.19121,
title = {CUPID: Curating Data your Robot Loves with Influence Functions},
author = {Christopher Agia and Rohan Sinha and Jingyun Yang and Rika Antonova and Marco Pavone and Haruki Nishimura and Masha Itkina and Jeannette Bohg},
journal= {arXiv preprint arXiv:2506.19121},
year = {2025}
}
备注
Project page: https://cupid-curation.github.io. 27 pages, 15 figures. Accepted to the Conference on Robot Learning (CoRL) 2025