English
Related papers

Related papers: FluentLip: A Phonemes-Based Two-stage Approach for…

200 papers

Recent studies have addressed intricate phonological phenomena in French, relying on either extensive linguistic knowledge or a significant amount of sentence-level pronunciation data. However, creating such resources is expensive and…

Computation and Language · Computer Science 2024-10-10 Hoyeon Lee , Hyeeun Jang , Jong-Hwan Kim , Jae-Min Kim

Existing dominant methods for audio generation include Generative Adversarial Networks (GANs) and diffusion-based methods like Flow Matching. GANs suffer from slow convergence during training, while diffusion methods require multi-step…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-10 Zengwei Yao , Wei Kang , Han Zhu , Liyong Guo , Lingxuan Ye , Fangjun Kuang , Weiji Zhuang , Zhaoqing Li , Zhifeng Han , Long Lin , Daniel Povey

We introduce SyncLipMAE, a self-supervised pretraining framework for talking-face video that learns synchronization-aware and transferable facial dynamics from unlabeled audio-visual streams. Our approach couples masked visual modeling with…

Artificial Intelligence · Computer Science 2026-01-07 Zeyu Ling , Xiaodong Gu , Jiangnan Tang , Changqing Zou

Lip-to-Speech (Lip2Speech) synthesis, which predicts corresponding speech from talking face images, has witnessed significant progress with various models and training strategies in a series of independent studies. However, existing studies…

Multimedia · Computer Science 2023-05-25 Zheng-Yan Sheng , Yang Ai , Zhen-Hua Ling

The goal of this paper is to synthesise talking faces with controllable facial motions. To achieve this goal, we propose two key ideas. The first is to establish a canonical space where every face has the same motion patterns but different…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Youngjoon Jang , Kyeongha Rho , Jong-Bin Woo , Hyeongkeun Lee , Jihwan Park , Youshin Lim , Byeong-Yeol Kim , Joon Son Chung

Silent speech interface is a promising technology that enables private communications in natural language. However, previous approaches only support a small and inflexible vocabulary, which leads to limited expressiveness. We leverage…

Human-Computer Interaction · Computer Science 2023-03-07 Zixiong Su , Shitao Fang , Jun Rekimoto

We present a novel one-shot talking head synthesis method that achieves disentangled and fine-grained control over lip motion, eye gaze&blink, head pose, and emotional expression. We represent different motions via disentangled latent…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Duomin Wang , Yu Deng , Zixin Yin , Heung-Yeung Shum , Baoyuan Wang

The goal of this paper is to learn strong lip reading models that can recognise speech in silent videos. Most prior works deal with the open-set visual speech recognition problem by adapting existing automatic speech recognition techniques…

Computer Vision and Pattern Recognition · Computer Science 2021-12-06 K R Prajwal , Triantafyllos Afouras , Andrew Zisserman

In this paper, we propose a neural end-to-end system for voice preserving, lip-synchronous translation of videos. The system is designed to combine multiple component models and produces a video of the original speaker speaking in the…

Recent developments in voice cloning and talking head generation demonstrate impressive capabilities in synthesizing natural speech and realistic lip synchronization. Current methods typically require and are trained on large scale datasets…

Sound · Computer Science 2025-09-17 Javeria Amir , Farwa Attaria , Mah Jabeen , Umara Noor , Zahid Rashid

Human lip-reading is a challenging task. It requires not only knowledge of underlying language but also visual clues to predict spoken words. Experts need certain level of experience and understanding of visual expressions learning to…

Computer Vision and Pattern Recognition · Computer Science 2018-02-16 M Faisal , Sanaullah Manzoor

Latent diffusion models have made great strides in generating expressive portrait videos with accurate lip-sync and natural motion from a single reference image and audio input. However, these models are far from real-time, often requiring…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Hanzhong Guo , Hongwei Yi , Daquan Zhou , Alexander William Bergman , Michael Lingelbach , Yizhou Yu

Despite the advancement in the domain of audio and audio-visual speech recognition, visual speech recognition systems are still quite under-explored due to the visual ambiguity of some phonemes. In this work, we propose a new lip-reading…

Computer Vision and Pattern Recognition · Computer Science 2021-08-10 Shahd Elashmawy , Marian Ramsis , Hesham M. Eraqi , Farah Eldeshnawy , Hadeel Mabrouk , Omar Abugabal , Nourhan Sakr

This study delves into the intricacies of synchronizing facial dynamics with multilingual audio inputs, focusing on the creation of visually compelling, time-synchronized animations through diffusion-based techniques. Diverging from…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Rui Zhang , Yixiao Fang , Zhengnan Lu , Pei Cheng , Zebiao Huang , Bin Fu

Visual cues, like lip motion, have been shown to improve the performance of Automatic Speech Recognition (ASR) systems in noisy environments. We propose LipGER (Lip Motion aided Generative Error Correction), a novel framework for leveraging…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-10 Sreyan Ghosh , Sonal Kumar , Ashish Seth , Purva Chiniya , Utkarsh Tyagi , Ramani Duraiswami , Dinesh Manocha

Automatic lip-reading (ALR) aims to automatically transcribe spoken content from a speaker's silent lip motion captured in video. Current mainstream lip-reading approaches only use a single visual encoder to model input videos of a single…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 He Wang , Pengcheng Guo , Xucheng Wan , Huan Zhou , Lei Xie

The goal of this work is to synchronise audio and video of a talking face using deep neural network models. Existing works have trained networks on proxy tasks such as cross-modal similarity learning, and then computed similarities between…

Computer Vision and Pattern Recognition · Computer Science 2021-03-22 You Jin Kim , Hee Soo Heo , Soo-Whan Chung , Bong-Jin Lee

The lip is a dominant dynamic facial unit when a person is speaking. Detecting lip events is beneficial to speech analysis and support for the hearing impaired. This paper proposes a 3D lip event detection pipeline that automatically…

Computer Vision and Pattern Recognition · Computer Science 2021-11-19 Jie Zhang , Robert B. Fisher

Lip reading aims to predict spoken language by analyzing lip movements. Despite advancements in lip reading technologies, performance degrades when models are applied to unseen speakers due to their sensitivity to variations in visual…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Jeong Hun Yeo , Chae Won Kim , Hyunjun Kim , Hyeongseop Rha , Seunghee Han , Wen-Huang Cheng , Yong Man Ro

Lip synchronization and audio-visual editing have emerged as fundamental challenges in multimodal learning, underpinning a wide range of applications, including film production, virtual avatars, and telepresence. Despite recent progress,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Lixiang Lin , Siyuan Jin , Jinshan Zhang