EgoVLA:基于以太视觉-语言-动作模型的学习
机器人学
2025-07-21 v3 人工智能
计算机视觉与模式识别
机器学习
摘要
机器人仿真学习已带来显著的机器人操控进展,但要求机器人硬件参与过程的本质限制了数据的规模。本文探索了使用以太视频训练视觉-语言-动作(Vision-Language-Action, VLA)模型的可能性。使用人类视频的优势不仅在于其规模,更在于场景和任务的丰富性。我们训练的VLA模型能够预测人类手腕和手部动作,然后通过逆运动学和动作 retargeting 将人类动作转换为机器人动作。我们使用少量机器人操控示范进行微调以获得机器人策略,称其为 EgoVLA。我们提出了称为 Ego Humanoid Manipulation Benchmark 的仿真基准,设计了多样化的双手操控任务并提供示范。我们使用 Ego Humanoid Manipulation Benchmark 对 EgoVLA 进行微调和评估,结果显示相较于基线方法显著提升,并验证了人类数据的重要性。视频可在我们的网站上查阅:https://rchalyang.github.io/EgoVLA
引用
@article{arxiv.2507.12440,
title = {EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos},
author = {Ruihan Yang and Qinxi Yu and Yecheng Wu and Rui Yan and Borui Li and An-Chieh Cheng and Xueyan Zou and Yunhao Fang and Xuxin Cheng and Ri-Zhao Qiu and Hongxu Yin and Sifei Liu and Song Han and Yao Lu and Xiaolong Wang},
journal= {arXiv preprint arXiv:2507.12440},
year = {2025}
}
备注
More videos can be found on our website: https://rchalyang.github.io/EgoVLA