English

3DGesPolicy: Phoneme-Aware Holistic Co-Speech Gesture Generation Based on Action Control

Computer Vision and Pattern Recognition 2026-01-27 v1 Artificial Intelligence Machine Learning Multimedia Sound

Abstract

Generating holistic co-speech gestures that integrate full-body motion with facial expressions suffers from semantically incoherent coordination on body motion and spatially unstable meaningless movements due to existing part-decomposed or frame-level regression methods, We introduce 3DGesPolicy, a novel action-based framework that reformulates holistic gesture generation as a continuous trajectory control problem through diffusion policy from robotics. By modeling frame-to-frame variations as unified holistic actions, our method effectively learns inter-frame holistic gesture motion patterns and ensures both spatially and semantically coherent movement trajectories that adhere to realistic motion manifolds. To further bridge the gap in expressive alignment, we propose a Gesture-Audio-Phoneme (GAP) fusion module that can deeply integrate and refine multi-modal signals, ensuring structured and fine-grained alignment between speech semantics, body motion, and facial expressions. Extensive quantitative and qualitative experiments on the BEAT2 dataset demonstrate the effectiveness of our 3DGesPolicy across other state-of-the-art methods in generating natural, expressive, and highly speech-aligned holistic gestures.

Keywords

Cite

@article{arxiv.2601.18451,
  title  = {3DGesPolicy: Phoneme-Aware Holistic Co-Speech Gesture Generation Based on Action Control},
  author = {Xuanmeng Sha and Liyun Zhang and Tomohiro Mashita and Naoya Chiba and Yuki Uranishi},
  journal= {arXiv preprint arXiv:2601.18451},
  year   = {2026}
}

Comments

13 pages, 5 figures

R2 v1 2026-07-01T09:20:21.707Z