English
Related papers

Related papers: EmoDiffTalk:Emotion-aware Diffusion for Editable 3…

200 papers

Recent advances in diffusion models have made significant progress in digital human generation. However, most existing models still struggle to maintain 3D consistency, temporal coherence, and motion accuracy. A key reason for these…

Graphics · Computer Science 2025-03-21 Xuan Gao , Jingtao Zhou , Dongyu Liu , Yuqi Zhou , Juyong Zhang

Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attention due to their ability to synthesize high-fidelity and…

Sound · Computer Science 2024-12-12 Yifan Xie , Tao Feng , Xin Zhang , Xiangyang Luo , Zixuan Guo , Weijiang Yu , Heng Chang , Fei Ma , Fei Richard Yu

Speech-driven 3D face animation aims to generate realistic facial expressions that match the speech content and emotion. However, existing methods often neglect emotional facial expressions or fail to disentangle them from speech content.…

Computer Vision and Pattern Recognition · Computer Science 2023-08-28 Ziqiao Peng , Haoyu Wu , Zhenbo Song , Hao Xu , Xiangyu Zhu , Jun He , Hongyan Liu , Zhaoxin Fan

We present 3DiFACE, a novel method for personalized speech-driven 3D facial animation and editing. While existing methods deterministically predict facial animations from speech, they overlook the inherent one-to-many relationship between…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Balamurugan Thambiraja , Sadegh Aliakbarian , Darren Cosker , Justus Thies

Audio-driven emotional 3D facial animation aims to generate synchronized lip movements and vivid facial expressions. However, most existing approaches focus on static and predefined emotion labels, limiting their diversity and naturalness.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Chang Liu , Ye Pan , Chenyang Ding , Susanto Rahardja , Xiaokang Yang

Generating animatable and editable 3D head avatars is essential for various applications in computer vision and graphics. Traditional 3D-aware generative adversarial networks (GANs), often using implicit fields like Neural Radiance Fields…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Guohao Li , Hongyu Yang , Yifang Men , Di Huang , Weixin Li , Ruijie Yang , Yunhong Wang

Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing methods can produce synchronized lip motions, they often rely on predefined identity or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Xuangeng Chu , Yuan Gan , Ziteng Cui , Shuhong Liu , Jian Wang , Bing Zhou , Tatsuya Harada

Speech-driven talking head generation is a critical yet challenging task with applications in augmented reality and virtual human modeling. While recent approaches using autoregressive and diffusion-based models have achieved notable…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Yihong Lin , Zhaoxin Fan , Xianjia Wu , Lingyu Xiong , Liang Peng , Xiandong Li , Wenxiong Kang , Songju Lei , Huang Xu

Achieving disentangled control over multiple facial motions and accommodating diverse input modalities greatly enhances the application and entertainment of the talking head generation. This necessitates a deep exploration of the decoupling…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Shuai Tan , Bin Ji , Mengxiao Bi , Ye Pan

Recent advances in Talking Head Generation (THG) have achieved impressive lip synchronization and visual quality through diffusion models; yet existing methods struggle to generate emotionally expressive portraits while preserving speaker…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Weipeng Tan , Chuming Lin , Chengming Xu , FeiFan Xu , Xiaobin Hu , Xiaozhong Ji , Junwei Zhu , Chengjie Wang , Yanwei Fu

Speech-driven 3D facial animation has been an attractive task in both academia and industry. Traditional methods mostly focus on learning a deterministic mapping from speech to animation. Recent approaches start to consider the…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Peng Chen , Xiaobao Wei , Ming Lu , Yitong Zhu , Naiming Yao , Xingyu Xiao , Hui Chen

Recent works on audio-driven talking head synthesis using Neural Radiance Fields (NeRF) have achieved impressive results. However, due to inadequate pose and expression control caused by NeRF implicit representation, these methods still…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Hongyun Yu , Zhan Qu , Qihang Yu , Jianchuan Chen , Zhonghua Jiang , Zhiwen Chen , Shengyu Zhang , Jimin Xu , Fei Wu , Chengfei Lv , Gang Yu

Audio-driven emotional 3D face animation aims to generate emotionally expressive talking heads with synchronized lip movements. However, previous research has often overlooked the influence of diverse emotions on facial expressions or…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Chang Liu , Qunfen Lin , Zijiao Zeng , Ye Pan

Recent advancements in 3D editing have highlighted the potential of text-driven methods in real-time, user-friendly AR/VR applications. However, current methods rely on 2D diffusion models without adequately considering multi-view…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Dong In Lee , Hyeongcheol Park , Jiyoung Seo , Eunbyung Park , Hyunje Park , Ha Dam Baek , Sangheon Shin , Sangmin Kim , Sangpil Kim

Multimodal-driven talking face generation refers to animating a portrait with the given pose, expression, and gaze transferred from the driving image and video, or estimated from the text and audio. However, existing methods ignore the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-10 Chao Xu , Shaoting Zhu , Junwei Zhu , Tianxin Huang , Jiangning Zhang , Ying Tai , Yong Liu

We introduce GaussianSpeech, a novel approach that synthesizes high-fidelity animation sequences of photo-realistic, personalized 3D human head avatars from spoken audio. To capture the expressive, detailed nature of human heads, including…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Shivangi Aneja , Artem Sevastopolsky , Tobias Kirschstein , Justus Thies , Angela Dai , Matthias Nießner

We introduce SEDTalker, an emotion-aware framework for speech-driven 3D facial animation that leverages frame-level speech emotion diarization to achieve fine-grained expressive control. Unlike prior approaches that rely on utterance-level…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Farzaneh Jafari , Stefano Berretti , Anup Basu

Implementing fine-grained emotion control is crucial for emotion generation tasks because it enhances the expressive capability of the generative model, allowing it to accurately and comprehensively capture and express various nuanced…

Computer Vision and Pattern Recognition · Computer Science 2024-02-05 Guanwen Feng , Haoran Cheng , Yunan Li , Zhiyuan Ma , Chaoneng Li , Zhihao Qian , Qiguang Miao , Chi-Man Pun

Speech emotion conversion is the task of converting the expressed emotion of a spoken utterance to a target emotion while preserving the lexical content and speaker identity. While most existing works in speech emotion conversion rely on…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-09 Navin Raj Prabhu , Bunlong Lay , Simon Welker , Nale Lehmann-Willenbrock , Timo Gerkmann

Recent advances in talking face generation have significantly improved facial animation synthesis. However, existing approaches face fundamental limitations: 3DMM-based methods maintain temporal consistency but lack fine-grained regional…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Kangwei Liu , Junwu Liu , Yun Cao , Jinlin Guo , Xiaowei Yi