English
Related papers

Related papers: Which phoneme-to-viseme maps best improve visual-o…

200 papers

Speech as a natural signal is composed of three parts - visemes (visual part of speech), phonemes (spoken part of speech), and language (the imposed structure). However, video as a medium for the delivery of speech and a multimedia…

Computation and Language · Computer Science 2020-06-17 Dhruva Sahrawat , Yaman Kumar , Shashwat Aggarwal , Yifang Yin , Rajiv Ratn Shah , Roger Zimmermann

This work describes an interactive decoding method to improve the performance of visual speech recognition systems using user input to compensate for the inherent ambiguity of the task. Unlike most phoneme-to-word decoding pipelines, which…

Computation and Language · Computer Science 2021-07-05 Brendan Shillingford , Yannis Assael , Misha Denil

Recent progress in Spoken Language Modeling has shown that learning language directly from speech is feasible. Generating speech through a pipeline that operates at the text level typically loses nuances, intonations, and non-verbal…

Computation and Language · Computer Science 2024-10-31 Maxime Poli , Emmanuel Chemla , Emmanuel Dupoux

During a conversation, our brain is responsible for combining information obtained from multiple senses in order to improve our ability to understand the message we are perceiving. Different studies have shown the importance of presenting…

Computer Vision and Pattern Recognition · Computer Science 2023-11-22 David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

In conventional speech recognition, phoneme-based models outperform grapheme-based models for non-phonetic languages such as English. The performance gap between the two typically reduces as the amount of training data is increased. In this…

Computation and Language · Computer Science 2019-09-25 Kazuki Irie , Rohit Prabhavalkar , Anjuli Kannan , Antoine Bruguier , David Rybach , Patrick Nguyen

Generating semantically coherent and visually accurate talking faces requires bridging the gap between linguistic meaning and facial articulation. Although audio-driven methods remain prevalent, their reliance on high-quality paired audio…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Xu Wang , Shengeng Tang , Fei Wang , Lechao Cheng , Dan Guo , Feng Xue , Richang Hong

Recent talking head synthesis works typically adopt speech features extracted from large-scale pre-trained acoustic models. However, the intrinsic many-to-many relationship between speech and lip motion causes phoneme-viseme alignment…

Graphics · Computer Science 2025-10-16 Yihuan Huang , Jiajun Liu , Yanzhen Ren , Jun Xue , Wuyang Liu , Zongkun Sun

Lipreading has a lot of potential applications such as in the domain of surveillance and video conferencing. Despite this, most of the work in building lipreading systems has been limited to classifying silent videos into classes…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-03 Yaman Kumar , Rohit Jain , Khwaja Mohd. Salik , Rajiv Ratn Shah , Yifang yin , Roger Zimmermann

Visual Automatic Speech Recognition (V-ASR) is a challenging task that involves interpreting spoken language solely from visual information, such as lip movements and facial expressions. This task is notably challenging due to the absence…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Matthew Kit Khinn Teng , Haibo Zhang , Takeshi Saitoh

In recent research, slight performance improvement is observed from automatic speech recognition systems to audio-visual speech recognition systems in the end-to-end framework with low-quality videos. Unmatching convergence rates and…

Computation and Language · Computer Science 2024-03-12 Yusheng Dai , Hang Chen , Jun Du , Xiaofei Ding , Ning Ding , Feijun Jiang , Chin-Hui Lee

Visual speech recognition is a challenging research problem with a particular practical application of aiding audio speech recognition in noisy scenarios. Multiple camera setups can be beneficial for the visual speech recognition systems in…

Computer Vision and Pattern Recognition · Computer Science 2018-06-29 Marina Zimmermann , Mostafa Mehdipour Ghazi , Hazım Kemal Ekenel , Jean-Philippe Thiran

We present a novel audio-driven facial animation approach that can generate realistic lip-synchronized 3D facial animations from the input audio. Our approach learns viseme dynamics from speech videos, produces animator-friendly viseme…

Graphics · Computer Science 2023-01-18 Linchao Bao , Haoxian Zhang , Yue Qian , Tangli Xue , Changhai Chen , Xuefei Zhe , Di Kang

In this paper, we address the problem of lip-voice synchronisation in videos containing human face and voice. Our approach is based on determining if the lips motion and the voice in a video are synchronised or not, depending on their…

Computer Vision and Pattern Recognition · Computer Science 2022-07-01 Venkatesh S. Kadandale , Juan F. Montesinos , Gloria Haro

In this paper, we defined the viseme (visual speech element) and described about the method of extracting visual feature vector. We defined the 10 visemes based on vowel by analyzing of Korean utterance and proposed the method of extracting…

Computation and Language · Computer Science 2014-11-19 Ha Jong Won , Li Gwang Chol , Kim Hyok Chol , Li Kum Song

Lip-reading aims to recognize speech content from videos via visual analysis of speakers' lip movements. This is a challenging task due to the existence of homophemes-words which involve identical or highly similar lip movements, as well as…

Computer Vision and Pattern Recognition · Computer Science 2019-09-04 Chenhao Wang

Numerous studies have investigated the effectiveness of audio-visual multimodal learning for speech enhancement (AVSE) tasks, seeking a solution that uses visual data as auxiliary and complementary input to reduce the noise of noisy speech…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-02 Shang-Yi Chuang , Hsin-Min Wang , Yu Tsao

Automatic lipreading has major potential impact for speech recognition, supplementing and complementing the acoustic modality. Most attempts at lipreading have been performed on small vocabulary tasks, due to a shortfall of appropriate…

Image and Video Processing · Electrical Eng. & Systems 2018-05-31 George Sterpu , Naomi Harte

Realistic, high-fidelity 3D facial animations are crucial for expressive avatar systems in human-computer interaction and accessibility. Although prior methods show promising quality, their reliance on the mesh domain limits their ability…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Alexandre Symeonidis-Herzig , Özge Mercanoğlu Sincan , Richard Bowden

Driven by deep learning techniques and large-scale datasets, recent years have witnessed a paradigm shift in automatic lip reading. While the main thrust of Visual Speech Recognition (VSR) was improving accuracy of Audio Speech Recognition…

Computer Vision and Pattern Recognition · Computer Science 2021-10-18 Marzieh Oghbaie , Arian Sabaghi , Kooshan Hashemifard , Mohammad Akbari

Speech-driven talking face synthesis (TFS) focuses on generating lifelike facial animations from audio input. Current TFS models perform well in English but unsatisfactorily in non-English languages, producing wrong mouth shapes and rigid…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Zibo Su , Kun Wei , Jiahua Li , Xu Yang , Cheng Deng