English
Related papers

Related papers: Lite Audio-Visual Speech Enhancement

200 papers

Visual-Semantic Embedding (VSE) aims to learn an embedding space where related visual and semantic instances are close to each other. Recent VSE models tend to design complex structures to pool visual and semantic features into fixed-length…

Multimedia · Computer Science 2022-10-06 Zijian Zhang , Chang Shu , Ya Xiao , Yuan Shen , Di Zhu , Jing Xiao , Youxin Chen , Jey Han Lau , Qian Zhang , Zheng Lu

Speaker extraction seeks to extract the target speech in a multi-talker scenario given an auxiliary reference. Such reference can be auditory, i.e., a pre-recorded speech, visual, i.e., lip movements, or contextual, i.e., phonetic sequence.…

Computer Vision and Pattern Recognition · Computer Science 2022-10-13 Junjie Li , Meng Ge , Zexu Pan , Longbiao Wang , Jianwu Dang

Despite the excellent performance of vision-language pre-trained models (VLPs) on conventional VQA task, they still suffer from two problems: First, VLPs tend to rely on language biases in datasets and fail to generalize to…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Qingyi Si , Yuanxin Liu , Zheng Lin , Peng Fu , Weiping Wang

Audiovisual speech recognition (AVSR) combines acoustic and visual cues to improve transcription robustness under challenging conditions but remains out of reach for most under-resourced languages due to the lack of labeled video corpora…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-10 Pol Buitrago , Pol Gàlvez , Oriol Pareras , Javier Hernando

The large amount of audiovisual content being shared online today has drawn substantial attention to the prospect of audiovisual self-supervised learning. Recent works have focused on each of these modalities separately, while others have…

Machine Learning · Computer Science 2021-06-18 Pingchuan Ma , Rodrigo Mira , Stavros Petridis , Björn W. Schuller , Maja Pantic

Audio-visual target speaker extraction (AV-TSE) models primarily rely on target visual cues to isolate the target speaker's voice from others. We know that humans leverage linguistic knowledge, such as syntax and semantics, to support…

Sound · Computer Science 2025-06-17 Wenxuan Wu , Shuai Wang , Xixin Wu , Helen Meng , Haizhou Li

Low-Light Image Enhancement (LLIE) is crucial for improving both human perception and computer vision tasks. This paper addresses two challenges in zero-reference LLIE: obtaining perceptually 'good' images using the Contrastive…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Yuka Ogino , Takahiro Toizumi , Atsushi Ito

In recent years, Automatic Speech Recognition (ASR) technology has approached human-level performance on conversational speech under relatively clean listening conditions. In more demanding situations involving distant microphones,…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-15 George Sterpu , Naomi Harte

Audio-visual speech recognition (AVSR) typically improves recognition accuracy in noisy environments by integrating noise-immune visual cues with audio signals. Nevertheless, high-noise audio inputs are prone to introducing adverse…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-09 Linzhi Wu , Xingyu Zhang , Hao Yuan , Yakun Zhang , Changyan Zheng , Liang Xie , Tiejun Liu , Erwei Yin

The scarcity of labeled audio-visual datasets is a constraint for training superior audio-visual speaker diarization systems. To improve the performance of audio-visual speaker diarization, we leverage pre-trained supervised and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-08 Huan Zhao , Li Zhang , Yue Li , Yannan Wang , Hongji Wang , Wei Rao , Qing Wang , Lei Xie

Audio-visual target speaker extraction (AV-TSE) models primarily rely on visual cues from the target speaker. However, humans also leverage linguistic knowledge, such as syntactic constraints, next word prediction, and prior knowledge of…

Sound · Computer Science 2025-11-11 Wenxuan Wu , Shuai Wang , Xixin Wu , Helen Meng , Haizhou Li

Voice interfaces integral to the human-computer interaction systems can benefit from speech emotion recognition (SER) to customize responses based on user emotions. Since humans convey emotions through multi-modal audio-visual cues,…

Machine Learning · Computer Science 2025-07-02 Varsha Pendyala , Pedro Morgado , William Sethares

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate a target speaker's voice from multi-speaker environments by leveraging visual cues as guidance. However, the performance of AV-TSE systems heavily relies on the quality of…

Sound · Computer Science 2025-07-22 Junjie Li , Wenxuan Wu , Shuai Wang , Zexu Pan , Kong Aik Lee , Helen Meng , Haizhou Li

Although deep learning algorithms are widely used for improving speech enhancement (SE) performance, the performance remains limited under highly challenging conditions, such as unseen noise or noise signals having low signal-to-noise…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-10 Yu-Wen Chen , Kuo-Hsuan Hung , Shang-Yi Chuang , Jonathan Sherman , Xugang Lu , Yu Tsao

Our objective is an audio-visual model for separating a single speaker from a mixture of sounds such as other speakers and background noise. Moreover, we wish to hear the speaker even when the visual cues are temporarily absent due to…

Computer Vision and Pattern Recognition · Computer Science 2019-07-12 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Multimodal speech recognition aims to improve the performance of automatic speech recognition (ASR) systems by leveraging additional visual information that is usually associated to the audio input. While previous approaches make crucial…

Sound · Computer Science 2022-04-29 Dan Oneata , Horia Cucu

In this paper, we aim to generate clean speech frame by frame from a live video stream and a noisy audio stream without relying on future inputs. To this end, we propose RT-LA-VocE, which completely re-designs every component of LA-VocE, a…

Sound · Computer Science 2024-07-11 Honglie Chen , Rodrigo Mira , Stavros Petridis , Maja Pantic

Visual cues, like lip motion, have been shown to improve the performance of Automatic Speech Recognition (ASR) systems in noisy environments. We propose LipGER (Lip Motion aided Generative Error Correction), a novel framework for leveraging…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-10 Sreyan Ghosh , Sonal Kumar , Ashish Seth , Purva Chiniya , Utkarsh Tyagi , Ramani Duraiswami , Dinesh Manocha

Despite the rapid advance of automatic speech recognition (ASR) technologies, accurate recognition of cocktail party speech characterised by the interference from overlapping speakers, background noise and room reverberation remains a…

Sound · Computer Science 2022-04-11 Guinan Li , Jianwei Yu , Jiajun Deng , Xunying Liu , Helen Meng

Audio-visual speech enhancement system is regarded as one of promising solutions for isolating and enhancing speech of desired speaker. Typical methods focus on predicting clean speech spectrum via a naive convolution neural network based…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-01 Xinmeng Xu , Yang Wang , Jie Jia , Binbin Chen , Dejun Li