English

The DiffuseStyleGesture+ entry to the GENEA Challenge 2023

Human-Computer Interaction 2023-08-29 v1 Artificial Intelligence Multimedia

Abstract

In this paper, we introduce the DiffuseStyleGesture+, our solution for the Generation and Evaluation of Non-verbal Behavior for Embodied Agents (GENEA) Challenge 2023, which aims to foster the development of realistic, automated systems for generating conversational gestures. Participants are provided with a pre-processed dataset and their systems are evaluated through crowdsourced scoring. Our proposed model, DiffuseStyleGesture+, leverages a diffusion model to generate gestures automatically. It incorporates a variety of modalities, including audio, text, speaker ID, and seed gestures. These diverse modalities are mapped to a hidden space and processed by a modified diffusion model to produce the corresponding gesture for a given speech input. Upon evaluation, the DiffuseStyleGesture+ demonstrated performance on par with the top-tier models in the challenge, showing no significant differences with those models in human-likeness, appropriateness for the interlocutor, and achieving competitive performance with the best model on appropriateness for agent speech. This indicates that our model is competitive and effective in generating realistic and appropriate gestures for given speech. The code, pre-trained models, and demos are available at https://github.com/YoungSeng/DiffuseStyleGesture/tree/DiffuseStyleGesturePlus/BEAT-TWH-main.

Keywords

Cite

@article{arxiv.2308.13879,
  title  = {The DiffuseStyleGesture+ entry to the GENEA Challenge 2023},
  author = {Sicheng Yang and Haiwei Xue and Zhensong Zhang and Minglei Li and Zhiyong Wu and Xiaofei Wu and Songcen Xu and Zonghong Dai},
  journal= {arXiv preprint arXiv:2308.13879},
  year   = {2023}
}

Comments

7 pages, 8 figures, ICMI 2023

R2 v1 2026-06-28T12:05:03.487Z