中文

Diffusion Implicit Policy for Unpaired Scene-aware Motion Synthesis

计算机视觉与模式识别 2025-11-11 v2

摘要

Scene-aware motion synthesis 近年来因其众多应用而受到广泛关注。先前的方法高度依赖配对 motion-scene 数据,当仅在少数特定场景上训练时,难以推广到多样化的场景。因此,我们提出了一个统一框架,称为 Diffusion Implicit Policy (DIP),用于 scene-aware motion synthesis,无需配对 motion-scene 数据。在本文中,我们在训练期间将人类-场景 interaction 从 motion synthesis 中解耦,然后在推理阶段将基于 interaction 的隐式 policy 引入 motion diffusion。通过迭代 diffusion 去噪和隐式 policy 优化来推导出合成 motion,同时维持 motion 的自然性和 interaction 的合理性。对于长期 motion synthesis,我们引入了在关节旋转 power space 中的 motion blending。实验表明,我们的框架在更自然的 motion 以及更可信的 interaction 方面优于最新方法。这也表明了利用 DIP 进行更通用任务和灵活场景下的 motion synthesis 的可行性。代码将在 https://github.com/jingyugong/DIP 上公开。

关键词

引用

@article{arxiv.2412.02261,
  title  = {Diffusion Implicit Policy for Unpaired Scene-aware Motion Synthesis},
  author = {Jingyu Gong and Chong Zhang and Fengqi Liu and Ke Fan and Qianyu Zhou and Xin Tan and Zhizhong Zhang and Yuan Xie},
  journal= {arXiv preprint arXiv:2412.02261},
  year   = {2025}
}