English
Related papers

Related papers: WhisperNetV2: SlowFast Siamese Network For Lip-Bas…

200 papers

We present ObamaNet, the first architecture that generates both audio and synchronized photo-realistic lip-sync videos from any new text. Contrary to other published lip-sync approaches, ours is only composed of fully trainable neural…

Computer Vision and Pattern Recognition · Computer Science 2018-01-08 Rithesh Kumar , Jose Sotelo , Kundan Kumar , Alexandre de Brebisson , Yoshua Bengio

The goal of this paper is to learn strong lip reading models that can recognise speech in silent videos. Most prior works deal with the open-set visual speech recognition problem by adapting existing automatic speech recognition techniques…

Computer Vision and Pattern Recognition · Computer Science 2021-12-06 K R Prajwal , Triantafyllos Afouras , Andrew Zisserman

Lip-reading models have been significantly improved recently thanks to powerful deep learning architectures. However, most works focused on frontal or near frontal views of the mouth. As a consequence, lip-reading performance seriously…

Computer Vision and Pattern Recognition · Computer Science 2019-11-15 Shiyang Cheng , Pingchuan Ma , Georgios Tzimiropoulos , Stavros Petridis , Adrian Bulat , Jie Shen , Maja Pantic

Significant progress has been made in speaker dependent Lip-to-Speech synthesis, which aims to generate speech from silent videos of talking faces. Current state-of-the-art approaches primarily employ non-autoregressive sequence-to-sequence…

Sound · Computer Science 2023-07-06 Neha Sahipjohn , Neil Shah , Vishal Tambrahalli , Vineet Gandhi

In recent years, DeepFake technology has achieved unprecedented success in high-quality video synthesis, but these methods also pose potential and severe security threats to humanity. DeepFake can be bifurcated into entertainment…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Weifeng Liu , Tianyi She , Jiawei Liu , Boheng Li , Dongyu Yao , Ziyou Liang , Run Wang

The generation of emotional talking faces from a single portrait image remains a significant challenge. The simultaneous achievement of expressive emotional talking and accurate lip-sync is particularly difficult, as expressiveness is often…

Computer Vision and Pattern Recognition · Computer Science 2023-12-22 Chenxu Zhang , Chao Wang , Jianfeng Zhang , Hongyi Xu , Guoxian Song , You Xie , Linjie Luo , Yapeng Tian , Xiaohu Guo , Jiashi Feng

Talking face synthesis has been widely studied in either appearance-based or warping-based methods. Previous works mostly utilize single face image as a source, and generate novel facial animations by merging other person's facial features.…

Computer Vision and Pattern Recognition · Computer Science 2019-11-22 Kuangxiao Gu , Yuqian Zhou , Thomas Huang

Automatic emotion recognition plays a significant role in the process of human computer interaction and the design of Internet of Things (IOT) technologies. Yet, a common problem in emotion recognition systems lies in the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2020-06-05 Kexin Feng , Theodora Chaspari

Existing lip-sync deepfake detectors rely on pixel artifacts or audio-visual correspondence, and both fail under generator or language shift because the features they learn are tied to the training distribution. We take a different…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Hao Chen , Junnan Xu

Lip reading aims to predict spoken language by analyzing lip movements. Despite advancements in lip reading technologies, performance degrades when models are applied to unseen speakers due to their sensitivity to variations in visual…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Jeong Hun Yeo , Chae Won Kim , Hyunjun Kim , Hyeongseop Rha , Seunghee Han , Wen-Huang Cheng , Yong Man Ro

Face super-resolution (FSR) under limited computational budgets remains challenging. Existing methods often treat all facial pixels equally, leading to suboptimal resource allocation and degraded performance. CNNs are sensitive to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Siyu Xu , Wenjie Li , Guangwei Gao , Jian Yang , Guo-Jun Qi , Chia-Wen Lin

Audio-driven talking face generation has received growing interest, particularly for applications requiring expressive and natural human-avatar interaction. However, most existing emotion-aware methods rely on a single modality (either…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Phyo Thet Yee , Dimitrios Kollias , Sudeepta Mishra , Abhinav Dhall

Recent studies have demonstrated that incorporating auxiliary information, such as speaker voiceprint or visual cues, can substantially improve Speech Enhancement (SE) performance. However, single-channel methods often yield suboptimal…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-06 Chihyun Liu , Jiaxuan Fan , Mingtung Sun , Michael Anthony , Mingsian R. Bai , Yu Tsao

Talking face generation with great practical significance has attracted more attention in recent audio-visual studies. How to achieve accurate lip synchronization is a long-standing challenge to be further investigated. Motivated by xxx, in…

Computer Vision and Pattern Recognition · Computer Science 2022-04-12 Ganglai Wang , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha

Lip-to-Speech (Lip2Speech) synthesis, which predicts corresponding speech from talking face images, has witnessed significant progress with various models and training strategies in a series of independent studies. However, existing studies…

Multimedia · Computer Science 2023-05-25 Zheng-Yan Sheng , Yang Ai , Zhen-Hua Ling

Lip sync has emerged as a promising technique for generating mouth movements from audio signals. However, synthesizing a high-resolution and photorealistic virtual news anchor is still challenging. Lack of natural appearance, visual…

Computer Vision and Pattern Recognition · Computer Science 2021-05-06 Ruobing Zheng , Zhou Zhu , Bo Song , Changjiang Ji

Lipreading is an important technique for facilitating human-computer interaction in noisy environments. Our previously developed self-supervised learning method, AV2vec, which leverages multimodal self-distillation, has demonstrated…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-11 Jing-Xuan Zhang , Tingzhi Mao , Longjiang Guo , Jin Li , Lichen Zhang

Lipreading is understanding speech from observed lip movements. An observed series of lip motions is an ordered sequence of visual lip gestures. These gestures are commonly known, but as yet are not formally defined, as `visemes'. In this…

Image and Video Processing · Electrical Eng. & Systems 2019-09-17 Helen Bear , Richard Harvey

The objective of this work is the effective extraction of spatial and dynamic features for Continuous Sign Language Recognition (CSLR). To accomplish this, we utilise a two-pathway SlowFast network, where each pathway operates at distinct…

Computer Vision and Pattern Recognition · Computer Science 2023-09-22 Junseok Ahn , Youngjoon Jang , Joon Son Chung

Transformers and Mamba, initially invented for natural language processing, have inspired backbone architectures for visual recognition. Recent studies integrated Local Attention Transformers with Mamba to capture both local details and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Meng Lou , Yunxiang Fu , Yizhou Yu