English
Related papers

Related papers: EmoTalkingGaussian: Continuous Emotion-conditioned…

200 papers

Most current audio-driven facial animation research primarily focuses on generating videos with neutral emotions. While some studies have addressed the generation of facial videos driven by emotional audio, efficiently generating…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Chuhang Ma , Shuai Tan , Ye Pan , Jiaolong Yang , Xin Tong

We present a novel approach for synthesizing 3D talking heads with controllable emotion, featuring enhanced lip synchronization and rendering quality. Despite significant progress in the field, prior methods still suffer from multi-view…

Computer Vision and Pattern Recognition · Computer Science 2024-08-02 Qianyun He , Xinya Ji , Yicheng Gong , Yuanxun Lu , Zhengyu Diao , Linjia Huang , Yao Yao , Siyu Zhu , Zhan Ma , Songcen Xu , Xiaofei Wu , Zixiao Zhang , Xun Cao , Hao Zhu

Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attention due to their ability to synthesize high-fidelity and…

Sound · Computer Science 2024-12-12 Yifan Xie , Tao Feng , Xin Zhang , Xiangyang Luo , Zixuan Guo , Weijiang Yu , Heng Chang , Fei Ma , Fei Richard Yu

Recent works on audio-driven talking head synthesis using Neural Radiance Fields (NeRF) have achieved impressive results. However, due to inadequate pose and expression control caused by NeRF implicit representation, these methods still…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Hongyun Yu , Zhan Qu , Qihang Yu , Jianchuan Chen , Zhonghua Jiang , Zhiwen Chen , Shengyu Zhang , Jimin Xu , Fei Wu , Chengfei Lv , Gang Yu

Audio-driven 3D talking head synthesis has advanced rapidly with Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). By leveraging rich pre-trained priors, few-shot methods enable instant personalization from just a few seconds…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Haolan Xu , Keli Cheng , Lei Wang , Ning Bi , Xiaoming Liu

Achieving high synchronization in the synthesis of realistic, speech-driven talking head videos presents a significant challenge. A lifelike talking head requires synchronized coordination of subject identity, lip movements, facial…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Ziqiao Peng , Wentao Hu , Junyuan Ma , Xiangyu Zhu , Xiaomei Zhang , Hao Zhao , Hui Tian , Jun He , Hongyan Liu , Zhaoxin Fan

We propose GaussianTalker, a novel framework for real-time generation of pose-controllable talking heads. It leverages the fast rendering capabilities of 3D Gaussian Splatting (3DGS) while addressing the challenges of directly controlling…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Kyusun Cho , Joungbin Lee , Heeji Yoon , Yeobin Hong , Jaehoon Ko , Sangjun Ahn , Seungryong Kim

Radiance fields have demonstrated impressive performance in synthesizing lifelike 3D talking heads. However, due to the difficulty in fitting steep appearance changes, the prevailing paradigm that presents facial motions by directly…

Computer Vision and Pattern Recognition · Computer Science 2024-07-08 Jiahe Li , Jiawei Zhang , Xiao Bai , Jin Zheng , Xin Ning , Jun Zhou , Lin Gu

This paper presents EGSTalker, a real-time audio-driven talking head generation framework based on 3D Gaussian Splatting (3DGS). Designed to enhance both speed and visual fidelity, EGSTalker requires only 3-5 minutes of training video to…

Sound · Computer Science 2025-10-13 Tianheng Zhu , Yinfeng Yu , Liejun Wang , Fuchun Sun , Wendong Zheng

Recent photo-realistic 3D talking head via 3D Gaussian Splatting still has significant shortcoming in emotional expression manipulation, especially for fine-grained and expansive dynamics emotional editing using multi-modal control. This…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Chang Liu , Tianjiao Jing , Chengcheng Ma , Xuanqi Zhou , Zhengxuan Lian , Qin Jin , Hongliang Yuan , Shi-Sheng Huang

Creating high-quality, generalizable speech-driven 3D talking heads remains a persistent challenge. Previous methods achieve satisfactory results for fixed viewpoints and small-scale audio variations, but they struggle with large head…

Computer Vision and Pattern Recognition · Computer Science 2025-07-11 Wentao Hu , Shunkai Li , Ziqiao Peng , Haoxian Zhang , Fan Shi , Xiaoqiang Liu , Pengfei Wan , Di Zhang , Hui Tian

Audio-driven talking head generation is a core component of digital avatars, and 3D Gaussian Splatting has shown strong performance in real-time rendering of high-fidelity talking heads. However, achieving precise control over fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Shaoyang Xie , Xiaofeng Cong , Baosheng Yu , Zhipeng Gui , Jie Gui , Yuan Yan Tang , James Tin-Yau Kwok

Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing methods can produce synchronized lip motions, they often rely on predefined identity or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Xuangeng Chu , Yuan Gan , Ziteng Cui , Shuhong Liu , Jian Wang , Bing Zhou , Tatsuya Harada

Talking Head Generation aims at synthesizing natural-looking talking videos from speech and a single portrait image. Previous 3D talking head generation methods have relied on domain-specific heuristics such as warping-based facial motion…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Tong Shi , Melonie de Almeida , Daniela Ivanova , Nicolas Pugeault , Paul Henderson

The task of audio-driven portrait animation involves generating a talking head video using an identity image and an audio track of speech. While many existing approaches focus on lip synchronization and video quality, few tackle the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Jian Zhang , Weijian Mai , Zhijun Zhang

Audio-driven emotional 3D face animation aims to generate emotionally expressive talking heads with synchronized lip movements. However, previous research has often overlooked the influence of diverse emotions on facial expressions or…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Chang Liu , Qunfen Lin , Zijiao Zeng , Ye Pan

We introduce GaussianSpeech, a novel approach that synthesizes high-fidelity animation sequences of photo-realistic, personalized 3D human head avatars from spoken audio. To capture the expressive, detailed nature of human heads, including…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Shivangi Aneja , Artem Sevastopolsky , Tobias Kirschstein , Justus Thies , Angela Dai , Matthias Nießner

Several works have developed end-to-end pipelines for generating lip-synced talking faces with various real-world applications, such as teaching and language translation in videos. However, these prior works fail to create realistic-looking…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Sahil Goyal , Shagun Uppal , Sarthak Bhagat , Yi Yu , Yifang Yin , Rajiv Ratn Shah

Implementing fine-grained emotion control is crucial for emotion generation tasks because it enhances the expressive capability of the generative model, allowing it to accurately and comprehensively capture and express various nuanced…

Computer Vision and Pattern Recognition · Computer Science 2024-02-05 Guanwen Feng , Haoran Cheng , Yunan Li , Zhiyuan Ma , Chaoneng Li , Zhihao Qian , Qiguang Miao , Chi-Man Pun

We introduce GenSync, a novel framework for multi-identity lip-synced video synthesis using 3D Gaussian Splatting. Unlike most existing 3D methods that require training a new model for each identity , GenSync learns a unified network that…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Anushka Agarwal , Muhammad Yusuf Hassan , Talha Chafekar
‹ Prev 1 2 3 10 Next ›