English
Related papers

Related papers: FaceXHuBERT: Text-less Speech-driven E(X)pressive …

200 papers

3D facial animation has attracted considerable attention due to its extensive applications in the multimedia field. Audio-driven 3D facial animation has been widely explored with promising results. However, multi-modal 3D facial animation,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Sijing Wu , Yunhao Li , Yichao Yan , Huiyu Duan , Ziwei Liu , Guangtao Zhai

Audio-driven 3D facial animation has been widely explored, but achieving realistic, human-like performance is still unsolved. This is due to the lack of available 3D datasets, models, and standard evaluation metrics. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2019-05-09 Daniel Cudeiro , Timo Bolkart , Cassidy Laidlaw , Anurag Ranjan , Michael J. Black

We present READ Avatars, a 3D-based approach for generating 2D avatars that are driven by audio input with direct and granular control over the emotion. Previous methods are unable to achieve realistic animation due to the many-to-many…

Computer Vision and Pattern Recognition · Computer Science 2023-03-02 Jack Saunders , Vinay Namboodiri

Self-supervised approaches for speech representation learning are challenged by three unique problems: (1) there are multiple sound units in each input utterance, (2) there is no lexicon of input sound units during the pre-training phase,…

Computation and Language · Computer Science 2021-06-15 Wei-Ning Hsu , Benjamin Bolte , Yao-Hung Hubert Tsai , Kushal Lakhotia , Ruslan Salakhutdinov , Abdelrahman Mohamed

Although existing speech-driven talking face generation methods achieve significant progress, they are far from real-world application due to the avatar-specific training demand and unstable lip movements. To address the above issues, we…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Haiming Zhang , Zhihao Yuan , Chaoda Zheng , Xu Yan , Baoyuan Wang , Guanbin Li , Song Wu , Shuguang Cui , Zhen Li

Audio-driven 3D talking avatar generation is increasingly important in virtual communication, digital humans, and interactive media, where avatars must preserve identity, synchronize lip motion with speech, express emotion, and exhibit…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Zhongju Wang , Zhenhong Sun , Beier Wang , Yifu Wang , Daoyi Dong , Huadong Mo , Hongdong Li

Self-supervised speech representation learning methods like wav2vec 2.0 and Hidden-unit BERT (HuBERT) leverage unlabeled speech data for pre-training and offer good representations for numerous speech processing tasks. Despite the success…

Computation and Language · Computer Science 2022-04-29 Heng-Jui Chang , Shu-wen Yang , Hung-yi Lee

Speech-driven animation has gained significant traction in recent years, with current methods achieving near-photorealistic results. However, the field remains underexplored regarding non-verbal communication despite evidence demonstrating…

Computer Vision and Pattern Recognition · Computer Science 2023-08-31 Antoni Bigata Casademunt , Rodrigo Mira , Nikita Drobyshev , Konstantinos Vougioukas , Stavros Petridis , Maja Pantic

Most current audio-driven facial animation research primarily focuses on generating videos with neutral emotions. While some studies have addressed the generation of facial videos driven by emotional audio, efficiently generating…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Chuhang Ma , Shuai Tan , Ye Pan , Jiaolong Yang , Xin Tong

All previous methods for audio-driven talking head generation assume the input audio to be clean with a neutral tone. As we show empirically, one can easily break these systems by simply adding certain background noise to the utterance or…

Computer Vision and Pattern Recognition · Computer Science 2019-10-03 Gaurav Mittal , Baoyuan Wang

Vivid talking face generation holds immense potential applications across diverse multimedia domains, such as film and game production. While existing methods accurately synchronize lip movements with input audio, they typically ignore…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Jiadong Liang , Feng Lu

We introduce FaceTalk, a novel generative approach designed for synthesizing high-fidelity 3D motion sequences of talking human heads from input audio signal. To capture the expressive, detailed nature of human heads, including hair, ears,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Shivangi Aneja , Justus Thies , Angela Dai , Matthias Nießner

We present a novel audio-driven facial animation approach that can generate realistic lip-synchronized 3D facial animations from the input audio. Our approach learns viseme dynamics from speech videos, produces animator-friendly viseme…

Graphics · Computer Science 2023-01-18 Linchao Bao , Haoxian Zhang , Yue Qian , Tangli Xue , Changhai Chen , Xuefei Zhe , Di Kang

Video recordings of speech contain correlated audio and visual information, providing a strong signal for speech representation learning from the speaker's lip movements and the produced sound. We introduce Audio-Visual Hidden Unit BERT…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-15 Bowen Shi , Wei-Ning Hsu , Kushal Lakhotia , Abdelrahman Mohamed

Self-supervised models have had great success in learning speech representations that can generalize to various downstream tasks. However, most self-supervised models require a large amount of compute and multiple GPUs to train,…

Computation and Language · Computer Science 2024-09-02 Tzu-Quan Lin , Hung-yi Lee , Hao Tang

This paper introduces MauBERT, a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representation learning. We continue HuBERT pre-training with supervision based on a…

Computation and Language · Computer Science 2025-12-23 Angelo Ortiz Tandazo , Manel Khentout , Youssef Benchekroun , Thomas Hueber , Emmanuel Dupoux

Speech-driven 3D facial animation seeks to produce lifelike facial expressions that are synchronized with the speech content and its emotional nuances, finding applications in various multimedia fields. However, previous methods often…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Yixuan Zhang , Qing Chang , Yuxi Wang , Guang Chen , Zhaoxiang Zhang , Junran Peng

Audio-Driven Talking Face Generation aims at generating realistic videos of talking faces, focusing on accurate audio-lip synchronization without deteriorating any identity-related visual details. Recent state-of-the-art methods are based…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Hazım Kemal Ekenel , Alexander Waibel

Speech-driven facial animation involves using a speech signal to generate realistic videos of talking faces. Recent deep learning approaches to facial synthesis rely on extracting low-dimensional representations and concatenating them,…

Real-time speech-driven 3D facial animation has been attractive in academia and industry. Traditional methods mainly focus on learning a deterministic mapping from speech to animation. Recent approaches start to consider the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Peng Chen , Xiaobao Wei , Ming Lu , Hui Chen , Feng Tian