English
Related papers

Related papers: ObamaNet: Photo-realistic lip-sync from text

200 papers

Audio-visual (AV) lip biometrics is a promising authentication technique that leverages the benefits of both the audio and visual modalities in speech communication. Previous works have demonstrated the usefulness of AV lip biometrics.…

Multimedia · Computer Science 2021-04-27 Meng Liu , Longbiao Wang , Kong Aik Lee , Hanyi Zhang , Chang Zeng , Jianwu Dang

Emotionally talking head video generation aims to generate expressive portrait videos with accurate lip synchronization and emotional facial expressions. Current methods rely on simple emotional labels, leading to insufficient semantic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Yahui Li , Yinfeng Yu , Liejun Wang , Shengjie Shen

Talking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head…

Multimedia · Computer Science 2023-09-21 Songlin Yang , Wei Wang , Jun Ling , Bo Peng , Xu Tan , Jing Dong

The goal of this work is to simultaneously generate natural talking faces and speech outputs from text. We achieve this by integrating Talking Face Generation (TFG) and Text-to-Speech (TTS) systems into a unified framework. We address the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Youngjoon Jang , Ji-Hoon Kim , Junseok Ahn , Doyeop Kwak , Hong-Sun Yang , Yoon-Cheol Ju , Il-Hwan Kim , Byeong-Yeol Kim , Joon Son Chung

Audio-Driven Talking Face Generation aims at generating realistic videos of talking faces, focusing on accurate audio-lip synchronization without deteriorating any identity-related visual details. Recent state-of-the-art methods are based…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Hazım Kemal Ekenel , Alexander Waibel

In this paper, we propose a talking face generation method that takes an audio signal as input and a short target video clip as reference, and synthesizes a photo-realistic video of the target face with natural lip motions, head poses, and…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Chenxu Zhang , Yifan Zhao , Yifei Huang , Ming Zeng , Saifeng Ni , Madhukar Budagavi , Xiaohu Guo

Cross-modality generation is an emerging topic that aims to synthesize data in one modality based on information in a different modality. In this paper, we consider a task of such: given an arbitrary audio speech and one lip image of…

Computer Vision and Pattern Recognition · Computer Science 2018-05-23 Lele Chen , Zhiheng Li , Ross K. Maddox , Zhiyao Duan , Chenliang Xu

We present a method for generating a video of a talking face. The method takes as inputs: (i) still images of the target face, and (ii) an audio speech segment; and outputs a video of the target face lip synched with the audio. The method…

Computer Vision and Pattern Recognition · Computer Science 2017-07-19 Joon Son Chung , Amir Jamaludin , Andrew Zisserman

High-quality AI-powered video dubbing demands precise audio-lip synchronization, high-fidelity visual generation, and faithful preservation of identity and background. Most existing methods rely on a mask-based training strategy, where the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Xindi Zhang , Dechao Meng , Steven Xiao , Qi Wang , Peng Zhang , Bang Zhang

Lip-to-speech (L2S) synthesis, which reconstructs speech from visual cues, faces challenges in accuracy and naturalness due to limited supervision in capturing linguistic content, accents, and prosody. In this paper, we propose RESOUND, a…

Sound · Computer Science 2025-05-29 Long-Khanh Pham , Thanh V. T. Tran , Minh-Tan Pham , Van Nguyen

Current Non-Audible Murmur (NAM)-to-speech techniques rely on voice cloning to simulate ground-truth speech from paired whispers. However, the simulated speech often lacks intelligibility and fails to generalize well across different…

Sound · Computer Science 2025-01-24 Neil Shah , Shirish Karande , Vineet Gandhi

Audio-Visual Speech-to-Speech Translation typically prioritizes improving translation quality and naturalness. However, an equally critical aspect in audio-visual content is lip-synchrony-ensuring that the movements of the lips match the…

Speech-driven facial animation is the process which uses speech signals to automatically synthesize a talking character. The majority of work in this domain creates a mapping from audio features to visual features. This often requires…

Audio and Speech Processing · Electrical Eng. & Systems 2018-07-20 Konstantinos Vougioukas , Stavros Petridis , Maja Pantic

To the best of our knowledge, we first present a live system that generates personalized photorealistic talking-head animation only driven by audio signals at over 30 fps. Our system contains three stages. The first stage is a deep neural…

Graphics · Computer Science 2021-09-27 Yuanxun Lu , Jinxiang Chai , Xun Cao

Talking face generation is the challenging task of synthesizing a natural and realistic face that requires accurate synchronization with a given audio. Due to co-articulation, where an isolated phone is influenced by the preceding or…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Se Jin Park , Minsu Kim , Jeongsoo Choi , Yong Man Ro

Audio-visual multi-modal modeling has been demonstrated to be effective in many speech related tasks, such as speech recognition and speech enhancement. This paper introduces a new time-domain audio-visual architecture for target speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-24 Jian Wu , Yong Xu , Shi-Xiong Zhang , Lian-Wu Chen , Meng Yu , Lei Xie , Dong Yu

While accurate lip synchronization has been achieved for arbitrary-subject audio-driven talking face generation, the problem of how to efficiently drive the head pose remains. Previous methods rely on pre-estimated structural information…

Computer Vision and Pattern Recognition · Computer Science 2021-04-23 Hang Zhou , Yasheng Sun , Wayne Wu , Chen Change Loy , Xiaogang Wang , Ziwei Liu

A lip-syncing deepfake is a digitally manipulated video in which a person's lip movements are created convincingly using AI models to match altered or entirely new audio. Lip-syncing deepfakes are a dangerous type of deepfakes as the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Soumyya Kanti Datta , Shan Jia , Siwei Lyu

Text-to-Video generation, which utilizes the provided text prompt to generate high-quality videos, has drawn increasing attention and achieved great success due to the development of diffusion models recently. Existing methods mainly rely…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Zirui Pan , Xin Wang , Yipeng Zhang , Hong Chen , Kwan Man Cheng , Yaofei Wu , Wenwu Zhu

We introduce GenSync, a novel framework for multi-identity lip-synced video synthesis using 3D Gaussian Splatting. Unlike most existing 3D methods that require training a new model for each identity , GenSync learns a unified network that…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Anushka Agarwal , Muhammad Yusuf Hassan , Talha Chafekar