中文

基于潜在空间 GAN 的 LS-GAN:人类运动合成

计算机视觉与模式识别 2025-01-06 v1 人工智能

摘要

以文本输入为条件的人类运动合成近年来因其在 various domains 如 gaming、film production 和 virtual reality 中的潜在 applications 而受到广泛关注。Conditioned Motion synthesis 采用 text input 输出对应 text 的 3D motion。虽然 previous works 探索了使用 raw motion data 和 latent space representations 结合 diffusion models 的 motion synthesis,但 these approaches 常 suffer from high training 和 inference times。在本文中,我们引入 novel framework 使用 Generative Adversarial Networks (GANs) in latent space 以 enable faster training 和 inference,同时 achieving results comparable to state-of-the-art diffusion methods。我们在 HumanML3D、HumanAct12 benchmarks 上进行实验,展示 remarkably simple latent space 中的 GAN 达到 FID 为 0.482,FLOPs reduction 超过 91% compared to latent diffusion model。我们的 work 开辟了使用 latent space GANs 实现 efficient 和 high-quality motion synthesis 的 new possibilities。

关键词

引用

@article{arxiv.2501.01449,
  title  = {LS-GAN: Human Motion Synthesis with Latent-space GANs},
  author = {Avinash Amballa and Gayathri Akkinapalli and Vinitra Muralikrishnan},
  journal= {arXiv preprint arXiv:2501.01449},
  year   = {2025}
}

备注

6 pages