English
Related papers

Related papers: Self-Supervised Learning of Deviation in Latent Re…

200 papers

Co-speech gestures play a crucial role in the interactions between humans and embodied conversational agents (ECA). Recent deep learning methods enable the generation of realistic, natural co-speech gestures synchronized with speech, but…

Artificial Intelligence · Computer Science 2024-06-25 Teo Guichoux , Laure Soulier , Nicolas Obin , Catherine Pelachaud

We propose a real-time system for synthesizing gestures directly from speech. Our data-driven approach is based on Generative Adversarial Neural Networks to model the speech-gesture relationship. We utilize the large amount of speaker video…

Computer Vision and Pattern Recognition · Computer Science 2022-08-08 Manuel Rebol , Christian Gütl , Krzysztof Pietroszek

Co-speech gesture generation is crucial for automatic digital avatar animation. However, existing methods suffer from issues such as unstable training and temporal inconsistency, particularly in generating high-fidelity and comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2023-08-30 Longbin Ji , Pengfei Wei , Yi Ren , Jinglin Liu , Chen Zhang , Xiang Yin

Creating a virtual avatar with semantically coherent gestures that are aligned with speech is a challenging task. Existing gesture generation research mainly focused on generating rhythmic beat gestures, neglecting the semantic context of…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Lanmiao Liu , Esam Ghaleb , Aslı Özyürek , Zerrin Yumak

We propose a novel learned deep prior of body motion for 3D hand shape synthesis and estimation in the domain of conversational gestures. Our model builds upon the insight that body motion and hand gestures are strongly correlated in…

Computer Vision and Pattern Recognition · Computer Science 2021-04-08 Evonne Ng , Shiry Ginosar , Trevor Darrell , Hanbyul Joo

Self-supervised, multi-modal learning has been successful in holistic representation of complex scenarios. This can be useful to consolidate information from multiple modalities which have multiple, versatile uses. Its application in…

Computer Vision and Pattern Recognition · Computer Science 2020-11-03 Aniruddha Tamhane , Jie Ying Wu , Mathias Unberath

Deriving co-speech 3D gestures has seen tremendous progress in virtual avatar animation. Yet, the existing methods often produce stiff and unreasonable gestures with unseen human speech inputs due to the limited 3D speech-gesture data. In…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Xingqun Qi , Hengyuan Zhang , Yatian Wang , Jiahao Pan , Chen Liu , Peng Li , Xiaowei Chi , Mengfei Li , Wei Xue , Shanghang Zhang , Wenhan Luo , Qifeng Liu , Yike Guo

Generating full-body human gestures based on speech signals remains challenges on quality and speed. Existing approaches model different body regions such as body, legs and hands separately, which fail to capture the spatial interactions…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Pinxin Liu , Luchuan Song , Junhua Huang , Haiyang Liu , Chenliang Xu

Co-speech gesture generation aims to synthesize realistic body movements that are semantically coherent with speech and faithful to a user-specified gestural style. Existing VQ-VAE based co-speech gesture generation methods improve…

Graphics · Computer Science 2026-05-11 Junchuan Zhao , Qifan Liang , Ye Wang

Embodied agents, in the form of virtual agents or social robots, are rapidly becoming more widespread. In human-human interactions, humans use nonverbal behaviours to convey their attitudes, feelings, and intentions. Therefore, this…

Artificial Intelligence · Computer Science 2026-04-30 Carson Yu Liu , Gelareh Mohammadi , Yang Song , Wafa Johal

Along with the explosion of large language models, improvements in speech synthesis, advancements in hardware, and the evolution of computer graphics, the current bottleneck in creating digital humans lies in generating character movements…

Human-Computer Interaction · Computer Science 2026-01-30 Thanh Hoang-Minh

Sign language videos are an important medium for spreading and learning sign language. However, most existing human image synthesis methods produce sign language images with details that are distorted, blurred, or structurally incorrect.…

Computer Vision and Pattern Recognition · Computer Science 2024-09-26 Tongkai Shi , Lianyu Hu , Fanhua Shang , Jichao Feng , Peidong Liu , Wei Feng

Diffusion models can synthesize realistic co-speech video from audio for various applications, such as video creation and virtual agents. However, existing diffusion-based methods are slow due to numerous denoising steps and costly…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Beijia Lu , Ziyi Chen , Jing Xiao , Jun-Yan Zhu

In face-to-face interaction, we use multiple modalities, including speech and gestures, to communicate information and resolve references to objects. However, how representational co-speech gestures refer to objects remains understudied…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Esam Ghaleb , Bulat Khaertdinov , Aslı Özyürek , Raquel Fernández

We introduce a novel approach for high-resolution talking head generation from a single image and audio input. Prior methods using explicit face models, like 3D morphable models (3DMM) and facial landmarks, often fall short in generating…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Sejong Yang , Seoung Wug Oh , Yang Zhou , Seon Joo Kim

Hand gestures are a natural means of interaction in Augmented Reality and Virtual Reality (AR/VR) applications. Recently, there has been an increased focus on removing the dependence of accurate hand gesture recognition on complex sensor…

Computer Vision and Pattern Recognition · Computer Science 2019-12-09 Varun Jain , Shivam Aggarwal , Suril Mehta , Ramya Hebbalaguppe

The generation of co-speech gestures for digital humans is an emerging area in the field of virtual human creation. Prior research has made progress by using acoustic and semantic information as input and adopting classify method to…

Sound · Computer Science 2024-04-16 Fan Zhang , Naye Ji , Fuxing Gao , Siyuan Zhao , Zhaohan Wang , Shunman Li

We present a framework for generating full-bodied photorealistic avatars that gesture according to the conversational dynamics of a dyadic interaction. Given speech audio, we output multiple possibilities of gestural motion for an…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Evonne Ng , Javier Romero , Timur Bagautdinov , Shaojie Bai , Trevor Darrell , Angjoo Kanazawa , Alexander Richard

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this problem by proposing a…

Computer Vision and Pattern Recognition · Computer Science 2020-03-06 Ran Yi , Zipeng Ye , Juyong Zhang , Hujun Bao , Yong-Jin Liu

With read-aloud speech synthesis achieving high naturalness scores, there is a growing research interest in synthesising spontaneous speech. However, human spontaneous face-to-face conversation has both spoken and non-verbal aspects (here,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-15 Shivam Mehta , Siyang Wang , Simon Alexanderson , Jonas Beskow , Éva Székely , Gustav Eje Henter