扩散变换器策略
摘要
最近的大型视觉语言动作模型在多样化的机器人数据集上进行预训练,展示了针对新环境进行泛化的潜力。然而,这些方法通常通过小型动作头预测单个离散或连续动作,这限制了处理多样化动作空间的能力。相反,我们采用大型多模态扩散变换器建模连续动作序列,称为扩散变换器策略(Diffusion Transformer Policy),通过大型变换器模型而非小型动作头对动作块进行直接去噪。通过利用变换器的扩展能力,该方法能够有效地跨越大型多样化机器人数据集建模连续末端执行器动作,实现更好的泛化性能。广泛的实验表明,扩散变换器策略在 Maniskill2、Libero、Calvin 和 SimplerEnv 上有效果和泛化,以及在真实 Franka 机械臂上实现了更好的性能,相较于 OpenVLA 和 Octo 在 SimplerEnv 上的 Real-to-Sim 基准、真实 Franka 机械臂和 Libero 上表现更佳。 Specifically, without bells and whistles, the proposed approach achieves state-of-the-art performance with only a single third-view camera stream in the Calvin task ABC->D, improving the average number of tasks completed in a row of 5 to 3.6, and the pretraining stage significantly facilitates the success sequence length on the Calvin by over 1.2.
引用
@article{arxiv.2410.15959,
title = {Diffusion Transformer Policy},
author = {Zhi Hou and Tianyi Zhang and Yuwen Xiong and Hengjun Pu and Chengyang Zhao and Ronglei Tong and Yu Qiao and Jifeng Dai and Yuntao Chen},
journal= {arXiv preprint arXiv:2410.15959},
year = {2025}
}
备注
preprint; New Project Page: https://robodita.github.io; revert unsuitable replacement