English
Related papers

Related papers: EAD-Net: Emotion-Aware Talking Head Generation wit…

200 papers

Several works have developed end-to-end pipelines for generating lip-synced talking faces with various real-world applications, such as teaching and language translation in videos. However, these prior works fail to create realistic-looking…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Sahil Goyal , Shagun Uppal , Sarthak Bhagat , Yi Yu , Yifang Yin , Rajiv Ratn Shah

Vivid talking face generation holds immense potential applications across diverse multimedia domains, such as film and game production. While existing methods accurately synchronize lip movements with input audio, they typically ignore…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Jiadong Liang , Feng Lu

Lipreading is the task of decoding text from the movement of a speaker's mouth. Traditional approaches separated the problem into two stages: designing or learning visual features, and prediction. More recent deep lipreading approaches are…

Machine Learning · Computer Science 2016-12-19 Yannis M. Assael , Brendan Shillingford , Shimon Whiteson , Nando de Freitas

Talking head synthesis, also known as speech-to-lip synthesis, reconstructs the facial motions that align with the given audio tracks. The synthesized videos are evaluated on mainly two aspects, lip-speech synchronization and image…

Machine Learning · Computer Science 2025-03-18 Xulin Fan , Heting Gao , Ziyi Chen , Peng Chang , Mei Han , Mark Hasegawa-Johnson

Although automatically animating audio-driven talking heads has recently received growing interest, previous efforts have mainly concentrated on achieving lip synchronization with the audio, neglecting two crucial elements for generating…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Shuai Tan , Bin Ji , Ye Pan

Talking face generation has gained significant attention as a core application of generative models. To enhance the expressiveness and realism of synthesized videos, emotion editing in talking face video plays a crucial role. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Chanhyuk Choi , Taesoo Kim , Donggyu Lee , Siyeol Jung , Taehwan Kim

Recent multi-modal video generation models have achieved high visual quality, but their prohibitive latency and limited temporal stability hinder real-time deployment. Streaming inference exacerbates these issues, leading to pronounced…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Rang Meng , Weipeng Wu , Yuming Li , Chenguang Ma

Recent advances in Talking Head Generation (THG) have achieved impressive lip synchronization and visual quality through diffusion models; yet existing methods struggle to generate emotionally expressive portraits while preserving speaker…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Weipeng Tan , Chuming Lin , Chengming Xu , FeiFan Xu , Xiaobin Hu , Xiaozhong Ji , Junwei Zhu , Chengjie Wang , Yanwei Fu

Audio-driven talking head generation is a significant and challenging task applicable to various fields such as virtual avatars, film production, and online conferences. However, the existing GAN-based models emphasize generating…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Jintao Tan , Xize Cheng , Lingyu Xiong , Lei Zhu , Xiandong Li , Xianjia Wu , Kai Gong , Minglei Li , Yi Cai

The objective of stylized speech-driven facial animation is to create animations that encapsulate specific emotional expressions. Existing methods often depend on pre-established emotional labels or facial expression templates, which may…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Yicheng Zhong , Huawei Wei , Peiji Yang , Zhisheng Wang

Audio-driven talking-head generation has advanced rapidly with diffusion-based generative models, yet producing temporally coherent videos with fine-grained motion control remains challenging. We propose DEMO, a flow-matching generative…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Peiyin Chen , Zhuowei Yang , Hui Feng , Sheng Jiang , Rui Yan

Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing methods can produce synchronized lip motions, they often rely on predefined identity or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Xuangeng Chu , Yuan Gan , Ziteng Cui , Shuhong Liu , Jian Wang , Bing Zhou , Tatsuya Harada

Impressive progress has been made in audio-driven 3D facial animation recently, but synthesizing 3D talking-head with rich emotion is still unsolved. This is due to the lack of 3D generative models and available 3D emotional dataset with…

Computer Vision and Pattern Recognition · Computer Science 2021-04-27 Qianyun Wang , Zhenfeng Fan , Shihong Xia

Audio-driven talking-head generation is a crucial and useful technology for virtual human interaction and film-making. While recent advances have focused on improving image fidelity and lip synchronization, generating accurate emotional…

Computer Vision and Pattern Recognition · Computer Science 2025-05-05 Wenqing Wang , Yun Fu

Audio-driven emotional 3D facial animation aims to generate synchronized lip movements and vivid facial expressions. However, most existing approaches focus on static and predefined emotion labels, limiting their diversity and naturalness.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Chang Liu , Ye Pan , Chenyang Ding , Susanto Rahardja , Xiaokang Yang

Animating still face images with deep generative models using a speech input signal is an active research topic and has seen important recent progress.However, much of the effort has been put into lip syncing and rendering quality while the…

Graphics · Computer Science 2024-12-13 Louis Airale , Dominique Vaufreydaz , Xavier Alameda-Pineda

End-to-end audio-conditioned latent diffusion models (LDMs) have been widely adopted for audio-driven portrait animation, demonstrating their effectiveness in generating lifelike and high-resolution talking videos. However, direct…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Chunyu Li , Chao Zhang , Weikai Xu , Jingyu Lin , Jinghui Xie , Weiguo Feng , Bingyue Peng , Cunjian Chen , Weiwei Xing

Talking face generation, also known as speech-to-lip generation, reconstructs facial motions concerning lips given coherent speech input. The previous studies revealed the importance of lip-speech synchronization and visual quality. Despite…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Jiadong Wang , Xinyuan Qian , Malu Zhang , Robby T. Tan , Haizhou Li

Recent advancements in audio-driven talking face generation have made great progress in lip synchronization. However, current methods often lack sufficient control over facial animation such as speaking style and emotional expression,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Baiqin Wang , Xiangyu Zhu , Fan Shen , Hao Xu , Zhen Lei

Diffusion models have recently advanced photorealistic human synthesis, although practical talking-head generation (THG) remains constrained by high inference latency, temporal instability such as flicker and identity drift, and imperfect…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Soumya Mazumdar , Vineet Kumar Rakesh