基于潜在空间 GAN 的 LS-GAN:人类运动合成
摘要
以文本输入为条件的人类运动合成近年来因其在 various domains 如 gaming、film production 和 virtual reality 中的潜在 applications 而受到广泛关注。Conditioned Motion synthesis 采用 text input 输出对应 text 的 3D motion。虽然 previous works 探索了使用 raw motion data 和 latent space representations 结合 diffusion models 的 motion synthesis,但 these approaches 常 suffer from high training 和 inference times。在本文中,我们引入 novel framework 使用 Generative Adversarial Networks (GANs) in latent space 以 enable faster training 和 inference,同时 achieving results comparable to state-of-the-art diffusion methods。我们在 HumanML3D、HumanAct12 benchmarks 上进行实验,展示 remarkably simple latent space 中的 GAN 达到 FID 为 0.482,FLOPs reduction 超过 91% compared to latent diffusion model。我们的 work 开辟了使用 latent space GANs 实现 efficient 和 high-quality motion synthesis 的 new possibilities。
引用
@article{arxiv.2501.01449,
title = {LS-GAN: Human Motion Synthesis with Latent-space GANs},
author = {Avinash Amballa and Gayathri Akkinapalli and Vinitra Muralikrishnan},
journal= {arXiv preprint arXiv:2501.01449},
year = {2025}
}
备注
6 pages