English

LaserHuman: Language-guided Scene-aware Human Motion Generation in Free Environment

Computer Vision and Pattern Recognition 2024-03-22 v2

Abstract

Language-guided scene-aware human motion generation has great significance for entertainment and robotics. In response to the limitations of existing datasets, we introduce LaserHuman, a pioneering dataset engineered to revolutionize Scene-Text-to-Motion research. LaserHuman stands out with its inclusion of genuine human motions within 3D environments, unbounded free-form natural language descriptions, a blend of indoor and outdoor scenarios, and dynamic, ever-changing scenes. Diverse modalities of capture data and rich annotations present great opportunities for the research of conditional motion generation, and can also facilitate the development of real-life applications. Moreover, to generate semantically consistent and physically plausible human motions, we propose a multi-conditional diffusion model, which is simple but effective, achieving state-of-the-art performance on existing datasets.

Keywords

Cite

@article{arxiv.2403.13307,
  title  = {LaserHuman: Language-guided Scene-aware Human Motion Generation in Free Environment},
  author = {Peishan Cong and Ziyi Wang and Zhiyang Dou and Yiming Ren and Wei Yin and Kai Cheng and Yujing Sun and Xiaoxiao Long and Xinge Zhu and Yuexin Ma},
  journal= {arXiv preprint arXiv:2403.13307},
  year   = {2024}
}
R2 v1 2026-06-28T15:26:51.088Z