English
Related papers

Related papers: SyncTalk: The Devil is in the Synchronization for …

200 papers

Co-speech gesture generation is to synthesize a gesture sequence that not only looks real but also matches with the input speech audio. Our method generates the movements of a complete upper body, including arms, hands, and the head.…

Computer Vision and Pattern Recognition · Computer Science 2021-11-30 Shenhan Qian , Zhi Tu , Yihao Zhi , Wen Liu , Shenghua Gao

We leverage the modern advancements in talking head generation to propose an end-to-end system for talking head video compression. Our algorithm transmits pivot frames intermittently while the rest of the talking head video is generated by…

Computer Vision and Pattern Recognition · Computer Science 2022-10-10 Madhav Agarwal , Anchit Gupta , Rudrabha Mukhopadhyay , Vinay P. Namboodiri , C V Jawahar

Speech-driven facial animation methods usually contain two main classes, 3D and 2D talking face, both of which attract considerable research attention in recent years. However, to the best of our knowledge, the research on 3D talking face…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Yixiang Zhuang , Baoping Cheng , Yao Cheng , Yuntao Jin , Renshuai Liu , Chengyang Li , Xuan Cheng , Jing Liao , Juncong Lin

Speech-driven facial animation is the process which uses speech signals to automatically synthesize a talking character. The majority of work in this domain creates a mapping from audio features to visual features. This often requires…

Audio and Speech Processing · Electrical Eng. & Systems 2018-07-20 Konstantinos Vougioukas , Stavros Petridis , Maja Pantic

For audio-driven visual dubbing, it remains a considerable challenge to uphold and highlight speaker's persona while synthesizing accurate lip synchronization. Existing methods fall short of capturing speaker's unique speaking style or…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Longhao Zhang , Shuang Liang , Zhipeng Ge , Tianshu Hu

Speech-driven 3D facial animation aims to synthesize realistic facial motion sequences from given audio, matching the speaker's speaking style. However, previous works often require priors such as class labels of a speaker or additional 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Hyung Kyu Kim , Sangmin Lee , Hak Gu Kim

Audio-driven talking head generation requires precise synchronization between facial animations and audio signals. This paper introduces ATL-Diff, a novel approach addressing synchronization limitations while reducing noise and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Hoang-Son Vo , Quang-Vinh Nguyen , Seungwon Kim , Hyung-Jeong Yang , Soonja Yeom , Soo-Hyung Kim

Audio-driven talking head generation has drawn much attention in recent years, and many efforts have been made in lip-sync, expressive facial expressions, natural head pose generation, and high video quality. However, no model has yet led…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Xusen Sun , Longhao Zhang , Hao Zhu , Peng Zhang , Bang Zhang , Xinya Ji , Kangneng Zhou , Daiheng Gao , Liefeng Bo , Xun Cao

High-quality, real-time talking head synthesis remains a fundamental challenge in computer vision. Existing reconstruction- and rendering-based methods typically rely on identity-specific models, limiting cross-identity generalization. To…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Peng Jia , Zhen Xiao , Jia Li , Xueliang Liu , Zhenzhen Hu , Lingyun Yu

The goal of this paper is to synthesise talking faces with controllable facial motions. To achieve this goal, we propose two key ideas. The first is to establish a canonical space where every face has the same motion patterns but different…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Youngjoon Jang , Kyeongha Rho , Jong-Bin Woo , Hyeongkeun Lee , Jihwan Park , Youshin Lim , Byeong-Yeol Kim , Joon Son Chung

Audio-driven talking head generation aims to create vivid and realistic videos from a static portrait and speech. Existing AR-based methods rely on intermediate facial representations, which limit their expressiveness and realism.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Yuzhe Weng , Haotian Wang , Yuanhong Yu , Jun Du , Shan He , Xiaoyan Wu , Haoran Xu

Emotional talking head generation has attracted growing attention. Previous methods, which are mainly GAN-based, still struggle to consistently produce satisfactory results across diverse emotions and cannot conveniently specify…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Yifeng Ma , Shiwei Zhang , Jiayu Wang , Xiang Wang , Yingya Zhang , Zhidong Deng

Estimating spoken content from silent videos is crucial for applications in Assistive Technology (AT) and Augmented Reality (AR). However, accurately mapping lip movement sequences in videos to words poses significant challenges due to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Junxiao Xue , Xiaozhen Liu , Xuecheng Wu , Fei Yu , Jun Wang

Speech-driven 3D facial animation has garnered lots of attention thanks to its broad range of applications. Despite recent advancements in achieving realistic lip motion, current methods fail to capture the nuanced emotional undertones…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Jisoo Kim , Jungbin Cho , Joonho Park , Soonmin Hwang , Da Eun Kim , Geon Kim , Youngjae Yu

End-to-end audio-conditioned latent diffusion models (LDMs) have been widely adopted for audio-driven portrait animation, demonstrating their effectiveness in generating lifelike and high-resolution talking videos. However, direct…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Chunyu Li , Chao Zhang , Weikai Xu , Jingyu Lin , Jinghui Xie , Weiguo Feng , Bingyue Peng , Cunjian Chen , Weiwei Xing

Recent advancements in speech-driven 3D talking head generation have made significant progress in lip synchronization. However, existing models still struggle to capture the perceptual alignment between varying speech characteristics and…

Graphics · Computer Science 2025-04-01 Lee Chae-Yeon , Oh Hyun-Bin , Han EunGi , Kim Sung-Bin , Suekyeong Nam , Tae-Hyun Oh

Most lip-to-speech (LTS) synthesis models are trained and evaluated under the assumption that the audio-video pairs in the dataset are perfectly synchronized. In this work, we show that the commonly used audio-visual datasets, such as GRID,…

Sound · Computer Science 2023-03-02 Zhe Niu , Brian Mak

Gaze and head movements play a central role in expressive 3D media, human-agent interaction, and immersive communication. Existing works often model facial components in isolation and lack mechanisms for generating personalized, style-aware…

Graphics · Computer Science 2026-01-05 Chengwei Shi , Chong Cao

Achieving disentangled control over multiple facial motions and accommodating diverse input modalities greatly enhances the application and entertainment of the talking head generation. This necessitates a deep exploration of the decoupling…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Shuai Tan , Bin Ji

Individuals have unique facial expression and head pose styles that reflect their personalized speaking styles. Existing one-shot talking head methods cannot capture such personalized characteristics and therefore fail to produce diverse…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Suzhen Wang , Yifeng Ma , Yu Ding , Zhipeng Hu , Changjie Fan , Tangjie Lv , Zhidong Deng , Xin Yu