English
Related papers

Related papers: Lip2Vec: Efficient and Robust Visual Speech Recogn…

200 papers

Visual Speech Recognition (VSR) is a task to predict a sentence or word from lip movements. Some works have been recently presented which use audio signals to supplement visual information. However, existing methods utilize only limited…

Computer Vision and Pattern Recognition · Computer Science 2023-05-09 Jeong Hun Yeo , Minsu Kim , Yong Man Ro

Recently audio-visual speech recognition (AVSR), which better leverages video modality as additional information to extend automatic speech recognition (ASR), has shown promising results in complex acoustic environments. However, there is…

Sound · Computer Science 2023-12-15 Fan Yu , Haoxu Wang , Ziyang Ma , Shiliang Zhang

Video-to-speech (V2S) synthesis, the task of generating speech directly from silent video input, is inherently more challenging than other speech synthesis tasks due to the need to accurately reconstruct both speech content and speaker…

Sound · Computer Science 2025-03-10 Yifan Liu , Yu Fang , Zhouhan Lin

Lip reading involves interpreting a speaker's speech by analyzing sequences of lip movements. Currently, most models regard the left and right halves of the lips as a symmetrical whole, lacking a thorough investigation of their differences.…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Zejun gu , Junxia jiang

The intuitive interaction between the audio and visual modalities is valuable for cross-modal self-supervised learning. This concept has been demonstrated for generic audiovisual tasks like video action recognition and acoustic scene…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-14 Abhinav Shukla , Stavros Petridis , Maja Pantic

Automatic speech recognition (ASR) has reached a level of accuracy in recent years, that even outperforms humans in transcribing speech to text. Nevertheless, all current ASR approaches show a certain weakness against ambient noise. To…

Sound · Computer Science 2023-12-22 Christopher Simic , Tobias Bocklet

Video recordings of speech contain correlated audio and visual information, providing a strong signal for speech representation learning from the speaker's lip movements and the produced sound. We introduce Audio-Visual Hidden Unit BERT…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-15 Bowen Shi , Wei-Ning Hsu , Kushal Lakhotia , Abdelrahman Mohamed

Visual speech recognition models traditionally consist of two stages, feature extraction and classification. Several deep learning approaches have been recently presented aiming to replace the feature extraction stage by automatically…

Computer Vision and Pattern Recognition · Computer Science 2019-07-10 Stavros Petridis , Yujiang Wang , Pingchuan Ma , Zuwei Li , Maja Pantic

State-of-the-art automatic speech recognition (ASR) systems perform well on healthy speech. However, the performance on impaired speech still remains an issue. The current study explores the usefulness of using Wav2Vec self-supervised…

Computation and Language · Computer Science 2022-04-05 Abner Hernandez , Paula Andrea Pérez-Toro , Elmar Nöth , Juan Rafael Orozco-Arroyave , Andreas Maier , Seung Hee Yang

Lip-reading is the operation of recognizing speech from lip movements. This is a difficult task because the movements of the lips when pronouncing the words are similar for some of them. Viseme is used to describe lip movements during a…

Computer Vision and Pattern Recognition · Computer Science 2021-11-09 Javad Peymanfard , Mohammad Reza Mohammadi , Hossein Zeinali , Nasser Mozayani

Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker's lip movements. This approach has been shown to yield improvements over audio-only…

For video-text retrieval, the use of CLIP has been a de facto choice. Since CLIP provides only image and text encoders, this consensus has led to a biased paradigm that entirely ignores the sound track of videos. While several attempts have…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Ruixiang Zhao , Zhihao Xu , Bangxiang Lan , Zijie Xin , Jingyu Liu , Xirong Li

Masked speech modeling (MSM) methods such as wav2vec2 or w2v-BERT learn representations over speech frames which are randomly masked within an utterance. While these methods improve performance of Automatic Speech Recognition (ASR) systems,…

This paper deals with Audio-Visual Speech Recognition (AVSR) under multimodal input corruption situations where audio inputs and visual inputs are both corrupted, which is not well addressed in previous research directions. Previous studies…

Multimedia · Computer Science 2023-03-21 Joanna Hong , Minsu Kim , Jeongsoo Choi , Yong Man Ro

The growing prevalence of online conferences and courses presents a new challenge in improving automatic speech recognition (ASR) with enriched textual information from video slides. In contrast to rare phrase lists, the slides within…

Sound · Computer Science 2024-01-15 Fan Yu , Haoxu Wang , Xian Shi , Shiliang Zhang

Recent advances in sophisticated synthetic speech generated from text-to-speech (TTS) or voice conversion (VC) systems cause threats to the existing automatic speaker verification (ASV) systems. Since such synthetic speech is generated from…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-15 Youngsik Eom , Yeonghyeon Lee , Ji Sub Um , Hoirin Kim

Speech is one of the most effective ways of communication among humans. Even though audio is the most common way of transmitting speech, very important information can be found in other modalities, such as vision. Vision is particularly…

Computation and Language · Computer Science 2016-11-22 Ramon Sanabria , Florian Metze , Fernando De La Torre

Vision is often used as a complementary modality for audio speech recognition (ASR), especially in the noisy environment where performance of solo audio modality significantly deteriorates. After combining visual modality, ASR is upgraded…

Computer Vision and Pattern Recognition · Computer Science 2020-05-14 Bo Xu , Cheng Lu , Yandong Guo , Jacob Wang

Self-supervision has recently shown great promise for learning visual and auditory speech representations from unlabelled data. In this work, we propose BRAVEn, an extension to the recent RAVEn method, which learns speech representations…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Alexandros Haliassos , Andreas Zinonos , Rodrigo Mira , Stavros Petridis , Maja Pantic

Audio-Visual Speech-to-Speech Translation typically prioritizes improving translation quality and naturalness. However, an equally critical aspect in audio-visual content is lip-synchrony-ensuring that the movements of the lips match the…