English
Related papers

Related papers: SyncAnyone: Implicit Disentanglement via Progressi…

200 papers

Audio-Driven Talking Face Generation aims at generating realistic videos of talking faces, focusing on accurate audio-lip synchronization without deteriorating any identity-related visual details. Recent state-of-the-art methods are based…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Hazım Kemal Ekenel , Alexander Waibel

In this paper, we present StyleLipSync, a style-based personalized lip-sync video generative model that can generate identity-agnostic lip-synchronizing video from arbitrary audio. To generate a video of arbitrary identities, we leverage…

Computer Vision and Pattern Recognition · Computer Science 2024-02-13 Taekyung Ki , Dongchan Min

End-to-end audio-conditioned latent diffusion models (LDMs) have been widely adopted for audio-driven portrait animation, demonstrating their effectiveness in generating lifelike and high-resolution talking videos. However, direct…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Chunyu Li , Chao Zhang , Weikai Xu , Jingyu Lin , Jinghui Xie , Weiguo Feng , Bingyue Peng , Cunjian Chen , Weiwei Xing

We present HighSync, an end-to-end diffusion-based framework for high-fidelity lip synchronization that generates photorealistic talking-face videos aligned with arbitrary input audio. Existing approaches consistently struggle to reconcile…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Saeed Firouzi Daghigh , Majid Iranpour Mobarekeh , Mostafa Alavi , Mehdi Bagheri

Talking head synthesis, also known as speech-to-lip synthesis, reconstructs the facial motions that align with the given audio tracks. The synthesized videos are evaluated on mainly two aspects, lip-speech synchronization and image…

Machine Learning · Computer Science 2025-03-18 Xulin Fan , Heting Gao , Ziyi Chen , Peng Chang , Mei Han , Mark Hasegawa-Johnson

Significant progress has been made in talking-face video generation research; however, precise lip-audio synchronization and high visual quality remain challenging in editing lip shapes based on input audio. This paper introduces JoyGen, a…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Qili Wang , Dajiang Wu , Zihang Xu , Junshi Huang , Jun Lv

Creating a realistic animatable avatar from a single static portrait remains challenging. Existing approaches often struggle to capture subtle facial expressions, the associated global body movements, and the dynamic background. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Mengchao Wang , Qiang Wang , Fan Jiang , Yaqi Fan , Yunpeng Zhang , Yonggang Qi , Kun Zhao , Mu Xu

We present VideoReTalking, a new system to edit the faces of a real-world talking head video according to input audio, producing a high-quality and lip-syncing output video even with a different emotion. Our system disentangles this…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Kun Cheng , Xiaodong Cun , Yong Zhang , Menghan Xia , Fei Yin , Mingrui Zhu , Xuan Wang , Jue Wang , Nannan Wang

Recent advances in diffusion-based lip-syncing generative models have demonstrated their ability to produce highly synchronized talking face videos for visual dubbing. Although these models excel at lip synchronization, they often struggle…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Yanyu Zhu , Lichen Bai , Jintao Xu , Hai-tao Zheng

Audio-driven visual dubbing aims to synchronize a video's lip movements with new speech but is fundamentally challenged by the lack of ideal training data: paired videos differing only in lip motion. Existing methods circumvent this via…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Xu He , Haoxian Zhang , Hejia Chen , Changyuan Zheng , Liyang Chen , Songlin Tang , Jiehui Huang , Xiaoqiang Liu , Pengfei Wan , Zhiyong Wu

A lip-syncing deepfake is a digitally manipulated video in which a person's lip movements are created convincingly using AI models to match altered or entirely new audio. Lip-syncing deepfakes are a dangerous type of deepfakes as the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Soumyya Kanti Datta , Shan Jia , Siwei Lyu

Lip sync is a fundamental audio-visual task. However, existing lip sync methods fall short of being robust in the wild. One important cause could be distracting factors on the visual input side, making extracting lip motion information…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Chun Wang

Audio-driven talking face generation aims to synthesize video with lip movements synchronized to input audio. However, current generative techniques face challenges in preserving intricate regional textures (skin, teeth). To address the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-06 Lingyu Xiong , Xize Cheng , Jintao Tan , Xianjia Wu , Xiandong Li , Lei Zhu , Fei Ma , Minglei Li , Huang Xu , Zhihu Hu

In this paper, we present a video-based learning framework for animating personalized 3D talking faces from audio. We introduce two training-time data normalizations that significantly improve data sample efficiency. First, we isolate and…

Computer Vision and Pattern Recognition · Computer Science 2021-06-09 Avisek Lahiri , Vivek Kwatra , Christian Frueh , John Lewis , Chris Bregler

In this paper, we propose a neural end-to-end system for voice preserving, lip-synchronous translation of videos. The system is designed to combine multiple component models and produces a video of the original speaker speaking in the…

Synthesizing personalized talking faces that uphold and highlight a speaker's unique style while maintaining lip-sync accuracy remains a significant challenge. A primary limitation of existing approaches is the intrinsic confounding of…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Renjie Lu , Xulong Zhang , Xiaoyang Qu , Jianzong Wang , Shangfei Wang

We study the problem of syncing the lip movement in a video with the audio stream. Our solution finds an optimal alignment using a dual-domain recurrent neural network that is trained on synthetic data we generate by dropping and…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Yoav Shalev , Lior Wolf

Combining face swapping with lip synchronization technology offers a cost-effective solution for customized talking face generation. However, directly cascading existing models together tends to introduce significant interference between…

Computer Vision and Pattern Recognition · Computer Science 2024-05-10 Zeren Zhang , Haibo Qin , Jiayu Huang , Yixin Li , Hui Lin , Yitao Duan , Jinwen Ma

For realistic talking head generation, creating natural head motion while maintaining accurate lip synchronization is essential. To fulfill this challenging task, we propose DisCoHead, a novel method to disentangle and control head pose and…

Computer Vision and Pattern Recognition · Computer Science 2023-03-15 Geumbyeol Hwang , Sunwon Hong , Seunghyun Lee , Sungwoo Park , Gyeongsu Chae

We present FlashLips, a two-stage, mask-free lip-sync system that decouples lips control from rendering and achieves real-time performance, with our U-Net variant running at over 100 FPS on a single GPU, while matching the visual quality of…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Andreas Zinonos , Michał Stypułkowski , Antoni Bigata , Stavros Petridis , Maja Pantic , Nikita Drobyshev