English
Related papers

Related papers: Towards Pose-invariant Lip-Reading

200 papers

3D pose transfer that aims to transfer the desired pose to a target mesh is one of the most challenging 3D generation tasks. Previous attempts rely on well-defined parametric human models or skeletal joints as driving pose sources. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-04 Haoyu Chen , Hao Tang , Ehsan Adeli , Guoying Zhao

Lipreading, the technology of decoding spoken content from silent videos of lip movements, holds significant application value in fields such as public security. However, due to the subtle nature of articulatory gestures, existing…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Matteo Rossi

In this paper, we propose a talking face generation method that takes an audio signal as input and a short target video clip as reference, and synthesizes a photo-realistic video of the target face with natural lip motions, head poses, and…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Chenxu Zhang , Yifan Zhao , Yifei Huang , Ming Zeng , Saifeng Ni , Madhukar Budagavi , Xiaohu Guo

We present LipDiffuser, a conditional diffusion model for lip-to-speech generation synthesizing natural and intelligible speech directly from silent video recordings. Our approach leverages the magnitude-preserving ablated diffusion model…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-27 Julius Richter , Danilo de Oliveira , Tal Peer , Timo Gerkmann

Realistic, high-fidelity 3D facial animations are crucial for expressive avatar systems in human-computer interaction and accessibility. Although prior methods show promising quality, their reliance on the mesh domain limits their ability…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Alexandre Symeonidis-Herzig , Özge Mercanoğlu Sincan , Richard Bowden

Multimodal speech recognition aims to improve the performance of automatic speech recognition (ASR) systems by leveraging additional visual information that is usually associated to the audio input. While previous approaches make crucial…

Sound · Computer Science 2022-04-29 Dan Oneata , Horia Cucu

This work focuses on the analysis that whether 3D face models can be learned from only the speech inputs of speakers. Previous works for cross-modal face synthesis study image generation from voices. However, image synthesis includes…

Graphics · Computer Science 2021-04-22 Cho-Ying Wu , Ke Xu , Chin-Cheng Hsu , Ulrich Neumann

We present a novel LSTM cell architecture capable of learning both intra- and inter-perspective relationships available in visual sequences captured from multiple perspectives. Our architecture adopts a novel recurrent joint learning…

Computer Vision and Pattern Recognition · Computer Science 2021-05-07 Alireza Sepas-Moghaddam , Fernando Pereira , Paulo Lobato Correia , Ali Etemad

This work presents a scalable solution to open-vocabulary visual speech recognition. To achieve this, we constructed the largest existing visual speech recognition dataset, consisting of pairs of text and video clips of faces speaking…

Existing face-swapping methods often deliver competitive results in constrained settings but exhibit substantial quality degradation when handling extreme facial poses. To improve facial pose robustness, explicit geometric features are…

Computer Vision and Pattern Recognition · Computer Science 2026-01-26 Jongmin Yu , Hyeontaek Oh , Zhongtian Sun , Angelica I Aviles-Rivero , Moongu Jeon , Jinhong Yang

Compared to facial expression recognition, expression synthesis requires a very high-dimensional mapping. This problem exacerbates with increasing image sizes and limits existing expression synthesis approaches to relatively small images.…

Computer Vision and Pattern Recognition · Computer Science 2020-11-19 Nazar Khan , Arbish Akram , Arif Mahmood , Sania Ashraf , Kashif Murtaza

Although there has been much progress in the area of facial expression recognition (FER), most existing methods suffer when presented with images that have been captured from viewing angles that are non-frontal and substantially different…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Shuvendu Roy , Ali Etemad

Lip-reading is to utilize the visual information of the speaker's lip movements to recognize words and sentences. Existing event-based lip-reading solutions integrate different frame rate branches to learn spatio-temporal features of…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Wenhao Zhang , Jun Wang , Yong Luo , Lei Yu , Wei Yu , Zheng He , Jialie Shen

Recognition of human poses and actions is crucial for autonomous systems to interact smoothly with people. However, cameras generally capture human poses in 2D as images and videos, which can have significant appearance variations across…

Computer Vision and Pattern Recognition · Computer Science 2021-11-19 Ting Liu , Jennifer J. Sun , Long Zhao , Jiaping Zhao , Liangzhe Yuan , Yuxiao Wang , Liang-Chieh Chen , Florian Schroff , Hartwig Adam

Facial pose estimation has gained a lot of attentions in many practical applications, such as human-robot interaction, gaze estimation and driver monitoring. Meanwhile, end-to-end deep learning-based facial pose estimation is becoming more…

Computer Vision and Pattern Recognition · Computer Science 2020-10-13 Zhaoxiang Liu , Zezhou Chen , Jinqiang Bai , Shaohua Li , Shiguo Lian

Recent methods for audio-driven talking head synthesis often optimize neural radiance fields (NeRF) on a monocular talking portrait video, leveraging its capability to render high-fidelity and 3D-consistent novel-view frames. However, they…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 Jaehoon Ko , Kyusun Cho , Joungbin Lee , Heeji Yoon , Sangmin Lee , Sangjun Ahn , Seungryong Kim

Deep learning has been impressively successful in the last decade in predicting human head poses from monocular images. However, for in-the-wild inputs the research community relies predominantly on a single training set, 300W-LP, of…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Michael Welter

Lipreading is an important technique for facilitating human-computer interaction in noisy environments. Our previously developed self-supervised learning method, AV2vec, which leverages multimodal self-distillation, has demonstrated…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-11 Jing-Xuan Zhang , Tingzhi Mao , Longjiang Guo , Jin Li , Lichen Zhang

The lip is a dominant dynamic facial unit when a person is speaking. Detecting lip events is beneficial to speech analysis and support for the hearing impaired. This paper proposes a 3D lip event detection pipeline that automatically…

Computer Vision and Pattern Recognition · Computer Science 2021-11-19 Jie Zhang , Robert B. Fisher

We present a new method for lightweight novel-view synthesis that generalizes to an arbitrary forward-facing scene. Recent approaches are computationally expensive, require per-scene optimization, or produce a memory-expensive…

Computer Vision and Pattern Recognition · Computer Science 2022-10-05 Pavel Solovev , Taras Khakhulin , Denis Korzhenkov
‹ Prev 1 8 9 10 Next ›