English

GTN-Bailando: Genre Consistent Long-Term 3D Dance Generation based on Pre-trained Genre Token Network

Sound 2023-04-26 v1 Multimedia Audio and Speech Processing

Abstract

Music-driven 3D dance generation has become an intensive research topic in recent years with great potential for real-world applications. Most existing methods lack the consideration of genre, which results in genre inconsistency in the generated dance movements. In addition, the correlation between the dance genre and the music has not been investigated. To address these issues, we propose a genre-consistent dance generation framework, GTN-Bailando. First, we propose the Genre Token Network (GTN), which infers the genre from music to enhance the genre consistency of long-term dance generation. Second, to improve the generalization capability of the model, the strategy of pre-training and fine-tuning is adopted.Experimental results on the AIST++ dataset show that the proposed dance generation framework outperforms state-of-the-art methods in terms of motion quality and genre consistency.

Keywords

Cite

@article{arxiv.2304.12704,
  title  = {GTN-Bailando: Genre Consistent Long-Term 3D Dance Generation based on Pre-trained Genre Token Network},
  author = {Haolin Zhuang and Shun Lei and Long Xiao and Weiqin Li and Liyang Chen and Sicheng Yang and Zhiyong Wu and Shiyin Kang and Helen Meng},
  journal= {arXiv preprint arXiv:2304.12704},
  year   = {2023}
}

Comments

Accepted by ICASSP2023.Demo page: https://im1eon.github.io/ICASSP23-GTNB-DG/