English

Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions

Computer Vision and Pattern Recognition 2025-04-07 v1

Abstract

We explore how body shapes influence human motion synthesis, an aspect often overlooked in existing text-to-motion generation methods due to the ease of learning a homogenized, canonical body shape. However, this homogenization can distort the natural correlations between different body shapes and their motion dynamics. Our method addresses this gap by generating body-shape-aware human motions from natural language prompts. We utilize a finite scalar quantization-based variational autoencoder (FSQ-VAE) to quantize motion into discrete tokens and then leverage continuous body shape information to de-quantize these tokens back into continuous, detailed motion. Additionally, we harness the capabilities of a pretrained language model to predict both continuous shape parameters and motion tokens, facilitating the synthesis of text-aligned motions and decoding them into shape-aware motions. We evaluate our method quantitatively and qualitatively, and also conduct a comprehensive perceptual study to demonstrate its efficacy in generating shape-aware motions.

Keywords

Cite

@article{arxiv.2504.03639,
  title  = {Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions},
  author = {Ting-Hsuan Liao and Yi Zhou and Yu Shen and Chun-Hao Paul Huang and Saayan Mitra and Jia-Bin Huang and Uttaran Bhattacharya},
  journal= {arXiv preprint arXiv:2504.03639},
  year   = {2025}
}

Comments

CVPR 2025. Project page: https://shape-move.github.io