English
Related papers

Related papers: DiffuseStyleGesture: Stylized Audio-Driven Co-Spee…

200 papers

Deriving co-speech 3D gestures has seen tremendous progress in virtual avatar animation. Yet, the existing methods often produce stiff and unreasonable gestures with unseen human speech inputs due to the limited 3D speech-gesture data. In…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Xingqun Qi , Hengyuan Zhang , Yatian Wang , Jiahao Pan , Chen Liu , Peng Li , Xiaowei Chi , Mengfei Li , Wei Xue , Shanghang Zhang , Wenhan Luo , Qifeng Liu , Yike Guo

Co-speech gesture generation is crucial for automatic digital avatar animation. However, existing methods suffer from issues such as unstable training and temporal inconsistency, particularly in generating high-fidelity and comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2023-08-30 Longbin Ji , Pengfei Wei , Yi Ren , Jinglin Liu , Chen Zhang , Xiang Yin

Existing methods for synthesizing 3D human gestures from speech have shown promising results, but they do not explicitly model the impact of emotions on the generated gestures. Instead, these methods directly output animations from speech…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Kiran Chhatre , Radek Daněček , Nikos Athanasiou , Giorgio Becherini , Christopher Peters , Michael J. Black , Timo Bolkart

Co-speech gesture generation is to synthesize a gesture sequence that not only looks real but also matches with the input speech audio. Our method generates the movements of a complete upper body, including arms, hands, and the head.…

Computer Vision and Pattern Recognition · Computer Science 2021-11-30 Shenhan Qian , Zhi Tu , Yihao Zhi , Wen Liu , Shenghua Gao

The generation of stylistic 3D facial animations driven by speech presents a significant challenge as it requires learning a many-to-many mapping between speech, style, and the corresponding natural facial motion. However, existing methods…

Computer Vision and Pattern Recognition · Computer Science 2024-05-15 Zhiyao Sun , Tian Lv , Sheng Ye , Matthieu Lin , Jenny Sheng , Yu-Hui Wen , Minjing Yu , Yong-Jin Liu

With read-aloud speech synthesis achieving high naturalness scores, there is a growing research interest in synthesising spontaneous speech. However, human spontaneous face-to-face conversation has both spoken and non-verbal aspects (here,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-15 Shivam Mehta , Siyang Wang , Simon Alexanderson , Jonas Beskow , Éva Székely , Gustav Eje Henter

Full-body gestures play a pivotal role in natural interactions and are crucial for achieving effective communication. Nevertheless, most existing studies primarily focus on the gesture generation of speakers, overlooking the vital role of…

Graphics · Computer Science 2025-05-09 Jinhe Huang , Yongkang Cheng , Yuming Hang , Gaoge Han , Jinewei Li , Jing Zhang , Xingjian Gu

While increasing attention has been paid to co-speech gesture synthesis, most previous works neglect to investigate hand gestures with explicit and essential semantics. In this paper, we study co-speech gesture generation with an emphasis…

Co-speech gesture generation is crucial for creating lifelike avatars and enhancing human-computer interactions by synchronizing gestures with speech. Despite recent advancements, existing methods struggle with accurately identifying the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Pinxin Liu , Pengfei Zhang , Hyeongwoo Kim , Pablo Garrido , Ari Shapiro , Kyle Olszewski

Generating co-speech gestures in real time requires both temporal coherence and efficient sampling. We introduce a novel framework for streaming gesture generation that extends Rolling Diffusion models with structured progressive noise…

Machine Learning · Computer Science 2025-11-20 Evgeniia Vu , Andrei Boiarov , Dmitry Vetrov

How can we teach robots or virtual assistants to gesture naturally? Can we go further and adapt the gesturing style to follow a specific speaker? Gestures that are naturally timed with corresponding speech during human communication are…

Computer Vision and Pattern Recognition · Computer Science 2020-07-27 Chaitanya Ahuja , Dong Won Lee , Yukiko I. Nakano , Louis-Philippe Morency

Gestures are non-verbal but important behaviors accompanying people's speech. While previous methods are able to generate speech rhythm-synchronized gestures, the semantic context of the speech is generally lacking in the gesticulations.…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Yihao Zhi , Xiaodong Cun , Xuelin Chen , Xi Shen , Wen Guo , Shaoli Huang , Shenghua Gao

Human motion modeling is important for many modern graphics applications, which typically require professional skills. In order to remove the skill barriers for laymen, recent motion generation methods can directly generate human motions…

Computer Vision and Pattern Recognition · Computer Science 2022-09-01 Mingyuan Zhang , Zhongang Cai , Liang Pan , Fangzhou Hong , Xinying Guo , Lei Yang , Ziwei Liu

The generation of co-speech gestures for digital humans is an emerging area in the field of virtual human creation. Prior research has made progress by using acoustic and semantic information as input and adopting classify method to…

Sound · Computer Science 2024-04-16 Fan Zhang , Naye Ji , Fuxing Gao , Siyuan Zhao , Zhaohan Wang , Shunman Li

The body movements accompanying speech aid speakers in expressing their ideas. Co-speech motion generation is one of the important approaches for synthesizing realistic avatars. Due to the intricate correspondence between speech and motion,…

Multimedia · Computer Science 2024-08-28 Sen Wang , Jiangning Zhang , Xin Tan , Zhifeng Xie , Chengjie Wang , Lizhuang Ma

While Diffusion Generative Models have achieved great success on image generation tasks, how to efficiently and effectively incorporate them into speech generation especially translation tasks remains a non-trivial problem. Specifically,…

Computation and Language · Computer Science 2023-10-27 Yongxin Zhu , Zhujin Gao , Xinyuan Zhou , Zhongyi Ye , Linli Xu

Co-speech gesture generation has significantly advanced human-computer interaction, yet speaker movements remain constrained due to the omission of text-driven non-spontaneous gestures (e.g., bowing while talking). Existing methods face two…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Fengyi Fang , Sicheng Yang , Wenming Yang

Recently, the multimedia community has witnessed the rise of diffusion models trained on large-scale multi-modal data for visual content creation, particularly in the field of text-to-image generation. In this paper, we propose a new task…

Computer Vision and Pattern Recognition · Computer Science 2023-11-10 Jingwen Chen , Yingwei Pan , Ting Yao , Tao Mei

Speech-driven gesture generation is an emerging field within virtual human creation. However, a significant challenge lies in accurately determining and processing the multitude of input features (such as acoustic, semantic, emotional,…

Sound · Computer Science 2024-03-19 Fan Zhang , Zhaohan Wang , Xin Lyu , Siyuan Zhao , Mengjian Li , Weidong Geng , Naye Ji , Hui Du , Fuxing Gao , Hao Wu , Shunman Li

Animating virtual characters with holistic co-speech gestures is a challenging but critical task. Previous systems have primarily focused on the weak correlation between audio and gestures, leading to physically unnatural outcomes that…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Yongkang Cheng , Shaoli Huang