English
Related papers

Related papers: Identity-Preserving Realistic Talking Face Generat…

200 papers

Speech-Driven Facial Animation (SDFA) has gained significant attention due to its applications in movies, video games, and virtual reality. However, most existing models are trained on single-language data, limiting their effectiveness in…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Federico Nocentini , Kwanggyoon Seo , Qingju Liu , Claudio Ferrari , Stefano Berretti , David Ferman , Hyeongwoo Kim , Pablo Garrido , Akin Caliskan

To improve the experiences of face-to-face conversation with avatar, this paper presents a novel conversation system. It is composed of two sequence-to-sequence models respectively for listening and speaking and a Generative Adversarial…

Computer Vision and Pattern Recognition · Computer Science 2019-08-22 Zezhou Chen , Zhaoxiang Liu , Huan Hu , Jinqiang Bai , Shiguo Lian , Fuyuan Shi , Kai Wang

The objective of stylized speech-driven facial animation is to create animations that encapsulate specific emotional expressions. Existing methods often depend on pre-established emotional labels or facial expression templates, which may…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Yicheng Zhong , Huawei Wei , Peiji Yang , Zhisheng Wang

Generating talking person portraits with arbitrary speech audio is a crucial problem in the field of digital human and metaverse. A modern talking face generation method is expected to achieve the goals of generalized audio-lip…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Zhenhui Ye , Jinzheng He , Ziyue Jiang , Rongjie Huang , Jiawei Huang , Jinglin Liu , Yi Ren , Xiang Yin , Zejun Ma , Zhou Zhao

The combination of highly realistic voice cloning, along with visually compelling avatar, face-swap, or lip-sync deepfake video generation, makes it relatively easy to create a video of anyone saying anything. Today, such deepfake…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Justin D. Norman , Hany Farid

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simultaneously achieve the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Zhenhui Ye , Tianyun Zhong , Yi Ren , Jiaqi Yang , Weichuang Li , Jiawei Huang , Ziyue Jiang , Jinzheng He , Rongjie Huang , Jinglin Liu , Chen Zhang , Xiang Yin , Zejun Ma , Zhou Zhao

We devise a cascade GAN approach to generate talking face video, which is robust to different face shapes, view angles, facial characteristics, and noisy audio conditions. Instead of learning a direct mapping from audio to video frames, we…

Computer Vision and Pattern Recognition · Computer Science 2019-05-13 Lele Chen , Ross K. Maddox , Zhiyao Duan , Chenliang Xu

Understanding the lip movement and inferring the speech from it is notoriously difficult for the common person. The task of accurate lip-reading gets help from various cues of the speaker and its contextual or environmental setting. Every…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Munender Varshney , Ravindra Yadav , Vinay P. Namboodiri , Rajesh M Hegde

In order to be widely applicable, speech-driven 3D head avatars must articulate their lips in accordance with speech, while also conveying the appropriate emotions with dynamically changing facial expressions. The key problem is that…

Graphics · Computer Science 2026-01-28 Radek Daněček , Carolin Schmitt , Senya Polikovsky , Michael J. Black

While deep learning technologies are now capable of generating realistic images confusing humans, the research efforts are turning to the synthesis of images for more concrete and application-specific purposes. Facial image generation based…

Computer Vision and Pattern Recognition · Computer Science 2020-06-11 Yeqi Bai , Tao Ma , Lipo Wang , Zhenjie Zhang

We introduce GaussianSpeech, a novel approach that synthesizes high-fidelity animation sequences of photo-realistic, personalized 3D human head avatars from spoken audio. To capture the expressive, detailed nature of human heads, including…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Shivangi Aneja , Artem Sevastopolsky , Tobias Kirschstein , Justus Thies , Angela Dai , Matthias Nießner

Facial recognition has become a widely used method for authentication and identification, with applications for secure access and locating missing persons. Its success is largely attributed to deep learning, which leverages large datasets…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Pedro Vidal , Bernardo Biesseck , Luiz E. L. Coelho , Roger Granada , David Menotti

We introduce SyncLipMAE, a self-supervised pretraining framework for talking-face video that learns synchronization-aware and transferable facial dynamics from unlabeled audio-visual streams. Our approach couples masked visual modeling with…

Artificial Intelligence · Computer Science 2026-01-07 Zeyu Ling , Xiaodong Gu , Jiangnan Tang , Changqing Zou

In the task of talking face generation, the objective is to generate a face video with lips synchronized to the corresponding audio while preserving visual details and identity information. Current methods face the challenge of learning…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Seymanur Aktı , Hazım Kemal Ekenel , Alexander Waibel

Audio-driven talking face generation has received growing interest, particularly for applications requiring expressive and natural human-avatar interaction. However, most existing emotion-aware methods rely on a single modality (either…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Phyo Thet Yee , Dimitrios Kollias , Sudeepta Mishra , Abhinav Dhall

Dynamic NeRFs have recently garnered growing attention for 3D talking portrait synthesis. Despite advances in rendering speed and visual quality, challenges persist in enhancing efficiency and effectiveness. We present R2-Talker, an…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Zhiling Ye , LiangGuo Zhang , Dingheng Zeng , Quan Lu , Ning Jiang

Lip motion accuracy is important for speech intelligibility, especially for users who are hard of hearing or second language learners. A high level of realism in lip movements is also required for the game and film production industries. 3D…

Graphics · Computer Science 2024-07-25 Rabab Algadhy , Yoshihiko Gotoh , Steve Maddock

In this work, we investigate the problem of lip-syncing a talking face video of an arbitrary identity to match a target speech segment. Current works excel at producing accurate lip movements on a static image or videos of specific people…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 K R Prajwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C V Jawahar

VR Facial Animation is necessary in applications requiring clear view of the face, even though a VR headset is worn. In our case, we aim to animate the face of an operator who is controlling our robotic avatar system. We propose a real-time…

Computer Vision and Pattern Recognition · Computer Science 2023-04-25 Andre Rochow , Max Schwarz , Michael Schreiber , Sven Behnke

Cross-modality generation is an emerging topic that aims to synthesize data in one modality based on information in a different modality. In this paper, we consider a task of such: given an arbitrary audio speech and one lip image of…

Computer Vision and Pattern Recognition · Computer Science 2018-05-23 Lele Chen , Zhiheng Li , Ross K. Maddox , Zhiyao Duan , Chenliang Xu
‹ Prev 1 8 9 10 Next ›