English
Related papers

Related papers: VideoReTalking: Audio-based Lip Synchronization fo…

200 papers

We introduce GenSync, a novel framework for multi-identity lip-synced video synthesis using 3D Gaussian Splatting. Unlike most existing 3D methods that require training a new model for each identity , GenSync learns a unified network that…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Anushka Agarwal , Muhammad Yusuf Hassan , Talha Chafekar

Audio-driven human animation methods, such as talking head and talking body generation, have made remarkable progress in generating synchronized facial movements and appealing visual quality videos. However, existing methods primarily focus…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Zhe Kong , Feng Gao , Yong Zhang , Zhuoliang Kang , Xiaoming Wei , Xunliang Cai , Guanying Chen , Wenhan Luo

Humans involuntarily tend to infer parts of the conversation from lip movements when the speech is absent or corrupted by external noise. In this work, we explore the task of lip to speech synthesis, i.e., learning to generate natural…

Computer Vision and Pattern Recognition · Computer Science 2020-05-19 K R Prajwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C V Jawahar

In this work, we propose a joint system combining a talking face generation system with a text-to-speech system that can generate multilingual talking face videos from only the text input. Our system can synthesize natural multilingual…

Computer Vision and Pattern Recognition · Computer Science 2022-05-16 Hyoung-Kyu Song , Sang Hoon Woo , Junhyeok Lee , Seungmin Yang , Hyunjae Cho , Youseong Lee , Dongho Choi , Kang-wook Kim

Lip-to-Speech (Lip2Speech) synthesis, which predicts corresponding speech from talking face images, has witnessed significant progress with various models and training strategies in a series of independent studies. However, existing studies…

Multimedia · Computer Science 2023-05-25 Zheng-Yan Sheng , Yang Ai , Zhen-Hua Ling

Creating realistic or stylized facial and lip sync animation is a tedious task. It requires lot of time and skills to sync the lips with audio and convey the right emotion to the character's face. To allow animators to spend more time on…

Graphics · Computer Science 2024-06-03 Bastien Arcelin , Nicolas Chaverou

We introduce FaceTalk, a novel generative approach designed for synthesizing high-fidelity 3D motion sequences of talking human heads from input audio signal. To capture the expressive, detailed nature of human heads, including hair, ears,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Shivangi Aneja , Justus Thies , Angela Dai , Matthias Nießner

The goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio. Unlike previous works that have focussed on recognising a limited number of words or phrases, we tackle lip reading as an…

Computer Vision and Pattern Recognition · Computer Science 2018-12-27 Triantafyllos Afouras , Joon Son Chung , Andrew Senior , Oriol Vinyals , Andrew Zisserman

The goal of this work is to reconstruct high quality speech from lip motions alone, a task also known as lip-to-speech. A key challenge of lip-to-speech systems is the one-to-many mapping caused by (1) the existence of homophenes and (2)…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-05 Ji-Hoon Kim , Jaehun Kim , Joon Son Chung

Individuals have unique facial expression and head pose styles that reflect their personalized speaking styles. Existing one-shot talking head methods cannot capture such personalized characteristics and therefore fail to produce diverse…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Suzhen Wang , Yifeng Ma , Yu Ding , Zhipeng Hu , Changjie Fan , Tangjie Lv , Zhidong Deng , Xin Yu

Direct speech-to-speech translation achieves high-quality results through the introduction of discrete units obtained from self-supervised learning. This approach circumvents delays and cascading errors associated with model cascading.…

Talking head synthesis is to synthesize a lip-synchronized talking head video using audio. Recently, the capability of NeRF to enhance the realism and texture details of synthesized talking heads has attracted the attention of researchers.…

Graphics · Computer Science 2025-02-21 Xiaoxing Liu , Zhilei Liu , Chongke Bi

Vision-guided speech generation aims to produce authentic speech from facial appearance or lip motions without relying on auditory signals, offering significant potential for applications such as dubbing in filmmaking and assisting…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Jiaxin Ye , Hongming Shan

Speech-driven 3D facial animation has been widely explored, with applications in gaming, character animation, virtual reality, and telepresence systems. State-of-the-art methods deform the face topology of the target actor to sync the input…

Computer Vision and Pattern Recognition · Computer Science 2023-01-03 Balamurugan Thambiraja , Ikhsanul Habibie , Sadegh Aliakbarian , Darren Cosker , Christian Theobalt , Justus Thies

Whenever we speak, our voice is accompanied by facial movements and expressions. Several recent works have shown the synthesis of highly photo-realistic videos of talking faces, but they either require a source video to drive the target…

Computer Vision and Pattern Recognition · Computer Science 2020-11-03 Prateek Manocha , Prithwijit Guha

Since facial actions such as lip movements contain significant information about speech content, it is not surprising that audio-visual speech enhancement methods are more accurate than their audio-only counterparts. Yet, state-of-the-art…

Computer Vision and Pattern Recognition · Computer Science 2022-04-01 Karren Yang , Dejan Markovic , Steven Krenn , Vasu Agrawal , Alexander Richard

We present a novel one-shot talking head synthesis method that achieves disentangled and fine-grained control over lip motion, eye gaze&blink, head pose, and emotional expression. We represent different motions via disentangled latent…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Duomin Wang , Yu Deng , Zixin Yin , Heung-Yeung Shum , Baoyuan Wang

Audio-driven talking head generation is advancing from 2D to 3D content. Notably, Neural Radiance Field (NeRF) is in the spotlight as a means to synthesize high-quality 3D talking head outputs. Unfortunately, this NeRF-based approach…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Gihoon Kim , Kwanggyoon Seo , Sihun Cha , Junyong Noh

The task of Visual Text-to-Speech (VisualTTS), also known as video dubbing, aims to generate speech synchronized with the lip movements in an input video, in additional to being consistent with the content of input text and cloning the…

Multimedia · Computer Science 2025-12-01 Yuyue Wang , Xin Cheng , Yihan Wu , Xihua Wang , Jinchuan Tian , Ruihua Song

Audio-driven 3D talking avatar generation is increasingly important in virtual communication, digital humans, and interactive media, where avatars must preserve identity, synchronize lip motion with speech, express emotion, and exhibit…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Zhongju Wang , Zhenhong Sun , Beier Wang , Yifu Wang , Daoyi Dong , Huadong Mo , Hongdong Li