OmniHumanoid:带有 Paired-Free 适配的流式跨体化视频生成
摘要
Cross-embodiment video generation 旨在跨不同 humanoid embodiment 传递运动,如 human-to-robot 和 robot-to-robot,实现 for embodied intelligence 的可扩展数据生成。这一设定下的主要挑战在于,运动动力学在 embodiment 之间部分可传递,而外观和形态学保持 embodiment-specific。现有方法常将这些因素entangle,且许多 require paired data for 每个 target embodiment,这限制了对新机器人的可扩展性。我们提出 OmniHumanoid,一个将可传递运动学习与 embodiment-specific 适配 factorize 的框架。该方法从跨多个 embodiment 的 motion-aligned paired videos 中学习共享运动 transfer model,同时仅通过 unpaired videos 结合 lightweight embodiment-specific adapters 适配新 embodiment。为减少运动 transfer 与 embodiment adaptation 之间的干扰,我们进一步引入 branch-isolated attention 设计,将 motion conditioning 从 embodiment-specific 调制中分离。此外,我们构建了一个包含 motion-aligned paired videos 的合成 cross-embodiment 数据集,涵盖多样 humanoid assets、场景和视角。在 synthetic 和 real-world benchmarks 上进行的实验表明,OmniHumanoid 在运动保真度和 embodiment consistency 方面实现了强劲表现,同时实现了对未知 humanoid embodiment 的可扩展适配,无需重新训练共享运动模型。
引用
@article{arxiv.2605.12038,
title = {OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation},
author = {Yiren Song and Xiyao Deng and Pei Yang and Yihan Wang and Mike Zheng Shou},
journal= {arXiv preprint arXiv:2605.12038},
year = {2026}
}