English
Related papers

Related papers: Looking Enhances Listening: Recovering Missing Spe…

200 papers

Augmented reality (AR) requires the seamless integration of visual, auditory, and linguistic channels for optimized human-computer interaction. While auditory and visual inputs facilitate real-time and contextual user guidance, the…

Computation and Language · Computer Science 2023-10-19 Jing Bi , Nguyen Manh Nguyen , Ali Vosoughi , Chenliang Xu

This work investigates retrieval augmented generation as an efficient strategy for automatic context discovery in context-aware Automatic Speech Recognition (ASR) system, in order to improve transcription accuracy in the presence of rare or…

Computation and Language · Computer Science 2025-11-20 Dimitrios Siskos , Stavros Papadopoulos , Pablo Peso Parada , Jisi Zhang , Karthikeyan Saravanan , Anastasios Drosou

Driven by large scale datasets and LLM based architectures, automatic speech recognition (ASR) systems have achieved remarkable improvements in accuracy. However, challenges persist for domain-specific terminology, and short utterances…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 Jinming Chen , Lu Wang , Zheshu Song , Wei Deng

With the surge of online meetings, it has become more critical than ever to provide high-quality speech audio and live captioning under various noise conditions. However, most monaural speech enhancement (SE) models introduce processing…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-08 Sefik Emre Eskimez , Xiaofei Wang , Min Tang , Hemin Yang , Zirun Zhu , Zhuo Chen , Huaming Wang , Takuya Yoshioka

Automatic transcriptions of consumer-generated multi-media content such as "Youtube" videos still exhibit high word error rates. Such data typically occupies a very broad domain, has been recorded in challenging conditions, with cheap…

Computation and Language · Computer Science 2017-12-08 Abhinav Gupta , Yajie Miao , Leonardo Neves , Florian Metze

Speech restoration aims at restoring full-band speech with high quality and intelligibility, considering a diverse set of distortions. MaskSR is a recently proposed generative model for this task. As other models of its kind, MaskSR attains…

Sound · Computer Science 2024-09-17 Xiaoyu Liu , Xu Li , Joan Serrà , Santiago Pascual

Human speech processing is inherently multimodal, where visual cues (lip movements) help to better understand the speech in noise. Lip-reading driven speech enhancement significantly outperforms benchmark audio-only approaches at low…

Computer Vision and Pattern Recognition · Computer Science 2019-09-24 Ahsan Adeel , Mandar Gogate , Amir Hussain

Recent work has shown that it is possible to train a single model to perform joint acoustic echo cancellation (AEC), speech enhancement, and voice separation, thereby serving as a unified frontend for robust automatic speech recognition…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-15 Tom O'Malley , Arun Narayanan , Quan Wang

In this work, we present SynesLM, an unified model which can perform three multimodal language understanding tasks: audio-visual automatic speech recognition(AV-ASR) and visual-aided speech/machine translation(VST/VMT). Unlike previous…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-02 Yichen Lu , Jiaqi Song , Xuankai Chang , Hengwei Bian , Soumi Maiti , Shinji Watanabe

Audiovisual speech recognition (AVSR) is a method to alleviate the adverse effect of noise in the acoustic signal. Leveraging recent developments in deep neural network-based speech recognition, we present an AVSR neural network…

Computer Vision and Pattern Recognition · Computer Science 2018-05-01 Michael Wand , Ngoc Thang Vu , Juergen Schmidhuber

In recent research, slight performance improvement is observed from automatic speech recognition systems to audio-visual speech recognition systems in the end-to-end framework with low-quality videos. Unmatching convergence rates and…

Computation and Language · Computer Science 2024-03-12 Yusheng Dai , Hang Chen , Jun Du , Xiaofei Ding , Ning Ding , Feijun Jiang , Chin-Hui Lee

During language acquisition, infants have the benefit of visual cues to ground spoken language. Robots similarly have access to audio and visual sensors. Recent work has shown that images and spoken captions can be mapped into a meaningful…

Computation and Language · Computer Science 2017-05-29 Herman Kamper , Shane Settle , Gregory Shakhnarovich , Karen Livescu

We present an approach to Audio-Visual Speech Recognition that builds on a pre-trained Whisper model. To infuse visual information into this audio-only model, we extend it with an AV fusion module and LoRa adapters, one of the most…

Sound · Computer Science 2025-02-05 Christopher Simic , Korbinian Riedhammer , Tobias Bocklet

Vision-language models can assess visual context in an image and generate descriptive text. While the generated text may be accurate and syntactically correct, it is often overly general. To address this, recent work has used optical…

Computer Vision and Pattern Recognition · Computer Science 2022-07-12 Wes Robbins , Zanyar Zohourianshahzadi , Jugal Kalita

Automatic Speech Recognition (ASR) has witnessed a profound research interest. Recent breakthroughs have given ASR systems different prospects such as faithfully transcribing spoken language, which is a pivotal advancement in building…

Computation and Language · Computer Science 2024-03-05 Ankitha Sudarshan , Vinay Samuel , Parth Patwa , Ibtihel Amara , Aman Chadha

We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues to improve the quality of the target speech signal. We show…

Text data is commonly utilized as a primary input to enhance Speech Emotion Recognition (SER) performance and reliability. However, the reliance on human-transcribed text in most studies impedes the development of practical SER systems,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-25 Yuanchao Li , Peter Bell , Catherine Lai

Automatic speech recognition (ASR) for dysarthric speech remains challenging due to data scarcity, particularly in non-English languages. To address this, we fine-tune a voice conversion model on English dysarthric speech (UASpeech) to…

Audio-visual speech enhancement (AVSE) methods use both audio and visual features for the task of speech enhancement and the use of visual features has been shown to be particularly effective in multi-speaker scenarios. In the majority of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-18 Shrishti Saha Shetu , Soumitro Chakrabarty , Emanuël A. P. Habets

Visual Automatic Speech Recognition (V-ASR) is a challenging task that involves interpreting spoken language solely from visual information, such as lip movements and facial expressions. This task is notably challenging due to the absence…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Matthew Kit Khinn Teng , Haibo Zhang , Takeshi Saitoh