English
Related papers

Related papers: CSTalk: Correlation Supervised Speech-driven 3D Em…

200 papers

This paper addresses the problem of generating lifelike holistic co-speech motions for 3D avatars, focusing on two key aspects: variability and coordination. Variability allows the avatar to exhibit a wide range of motions even with similar…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Yifei Liu , Qiong Cao , Yandong Wen , Huaiguang Jiang , Changxing Ding

In order to be widely applicable, speech-driven 3D head avatars must articulate their lips in accordance with speech, while also conveying the appropriate emotions with dynamically changing facial expressions. The key problem is that…

Graphics · Computer Science 2026-01-28 Radek Daněček , Carolin Schmitt , Senya Polikovsky , Michael J. Black

In this project, we aim to build a Text-to-Speech system able to produce speech with a controllable emotional expressiveness. We propose a methodology for solving this problem in three main steps. The first is the collection of emotional…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-08 Noé Tits

Talking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limitations: mesh-based 3D methods can model dual-person dialogue…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Peng Chen , Xiaobao Wei , Yi Yang , Naiming Yao , Hui Chen , Feng Tian

This paper presents FaceXHuBERT, a text-less speech-driven 3D facial animation generation method that allows to capture personalized and subtle cues in speech (e.g. identity, emotion and hesitation). It is also very robust to background…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Kazi Injamamul Haque , Zerrin Yumak

Cross-modality generation is an emerging topic that aims to synthesize data in one modality based on information in a different modality. In this paper, we consider a task of such: given an arbitrary audio speech and one lip image of…

Computer Vision and Pattern Recognition · Computer Science 2018-05-23 Lele Chen , Zhiheng Li , Ross K. Maddox , Zhiyao Duan , Chenliang Xu

The recent advances in deep learning have made it possible to generate photo-realistic images by using neural networks and even to extrapolate video frames from an input video clip. In this paper, for the sake of both furthering this…

Computer Vision and Pattern Recognition · Computer Science 2018-08-10 Lijie Fan , Wenbing Huang , Chuang Gan , Junzhou Huang , Boqing Gong

In this paper, we propose a novel audio-driven talking head method capable of simultaneously generating highly expressive facial expressions and hand gestures. Unlike existing methods that focus on generating full-body or half-body poses,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Linrui Tian , Siqi Hu , Qi Wang , Bang Zhang , Liefeng Bo

The synthesis of 3D facial animations from speech has garnered considerable attention. Due to the scarcity of high-quality 4D facial data and well-annotated abundant multi-modality labels, previous methods often suffer from limited realism…

Computer Vision and Pattern Recognition · Computer Science 2024-01-31 Qingcheng Zhao , Pengyu Long , Qixuan Zhang , Dafei Qin , Han Liang , Longwen Zhang , Yingliang Zhang , Jingyi Yu , Lan Xu

While previous audio-driven talking head generation (THG) methods generate head poses from driving audio, the generated poses or lips cannot match the audio well or are not editable. In this study, we propose \textbf{PoseTalk}, a THG system…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Jun Ling , Yiwen Wang , Han Xue , Rong Xie , Li Song

Talking face generation has gained significant attention as a core application of generative models. To enhance the expressiveness and realism of synthesized videos, emotion editing in talking face video plays a crucial role. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Chanhyuk Choi , Taesoo Kim , Donggyu Lee , Siyeol Jung , Taehwan Kim

Responsive and accurate facial expression recognition is crucial to human-robot interaction for daily service robots. Nowadays, event cameras are becoming more widely adopted as they surpass RGB cameras in capturing facial expression…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Zhe Wang , Qijin Song , Yucen Peng , Weibang Bai

Facial expressions are an ideal means of communicating one's emotions or intentions to others. This overview will focus on human facial expression recognition as well as robotic facial expression generation. In the case of human facial…

Robotics · Computer Science 2022-02-09 Niyati Rawal , Ruth Maria Stock-Homburg

In dyadic interaction, predicting the listener's facial reactions is challenging as different reactions could be appropriate in response to the same speaker's behaviour. Previous approaches predominantly treated this task as an…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Cheng Luo , Siyang Song , Weicheng Xie , Micol Spitale , Zongyuan Ge , Linlin Shen , Hatice Gunes

Audio-driven talking head generation is advancing from 2D to 3D content. Notably, Neural Radiance Field (NeRF) is in the spotlight as a means to synthesize high-quality 3D talking head outputs. Unfortunately, this NeRF-based approach…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Gihoon Kim , Kwanggyoon Seo , Sihun Cha , Junyong Noh

This paper proposes a talking face generation method named "CP-EB" that takes an audio signal as input and a person image as reference, to synthesize a photo-realistic people talking video with head poses controlled by a short video clip…

Computer Vision and Pattern Recognition · Computer Science 2023-11-16 Jianzong Wang , Yimin Deng , Ziqi Liang , Xulong Zhang , Ning Cheng , Jing Xiao

Recent advances in facial expression synthesis have shown promising results using diverse expression representations including facial action units. Facial action units for an elaborate facial expression synthesis need to be intuitively…

Computer Vision and Pattern Recognition · Computer Science 2020-07-17 Joanna Hong , Jung Uk Kim , Sangmin Lee , Yong Man Ro

Speech-driven talking head generation is a critical yet challenging task with applications in augmented reality and virtual human modeling. While recent approaches using autoregressive and diffusion-based models have achieved notable…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Yihong Lin , Zhaoxin Fan , Xianjia Wu , Lingyu Xiong , Liang Peng , Xiandong Li , Wenxiong Kang , Songju Lei , Huang Xu

Recent studies on 3D hand reconstruction have demonstrated the effectiveness of synthetic training data to improve estimation performance. However, most methods rely on game engines to synthesize hand images, which often lack diversity in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zhuoran Zhao , Xianghao Kong , Linlin Yang , Zheng Wei , Pan Hui , Anyi Rao

Generative models have surged in popularity recently due to their ability to produce high-quality images and video. However, steering these models to produce images with specific attributes and precise control remains challenging. Humans,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Tuomas Varanka , Huai-Qian Khor , Yante Li , Mengting Wei , Hanwei Kung , Nicu Sebe , Guoying Zhao