English
Related papers

Related papers: Real-Time Audio-Visual End-to-End Speech Enhanceme…

200 papers

As an indispensable part of modern human-computer interaction system, speech synthesis technology helps users get the output of intelligent machine more easily and intuitively, thus has attracted more and more attention. Due to the…

Sound · Computer Science 2021-04-21 Zhaoxi Mu , Xinyu Yang , Yizhuo Dong

This paper presents a complete hardware and software pipeline for real-time speech enhancement in noisy and reverberant conditions. The device consists of a microphone array and a camera mounted on eyeglasses, connected to an embedded…

Speech activity detection (SAD) plays an important role in current speech processing systems, including automatic speech recognition (ASR). SAD is particularly difficult in environments with acoustic noise. A practical solution is to…

Computation and Language · Computer Science 2023-05-15 Fei Tao , Carlos Busso

Recently, an audio-visual speech generative model based on variational autoencoder (VAE) has been proposed, which is combined with a nonnegative matrix factorization (NMF) model for noise variance to perform unsupervised speech enhancement.…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-12 Mostafa Sadeghi , Xavier Alameda-Pineda

We present a live demonstration for RAVEN, a real-time audio-visual speech enhancement system designed to run entirely on a CPU. In single-channel, audio-only settings, speech enhancement is traditionally approached as the task of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-26 T. Aleksandra Ma , Sile Yin , Li-Chia Yang , Shuo Zhang

This work proposes an efficient method to enhance the quality of corrupted speech signals by leveraging both acoustic and visual cues. While existing diffusion-based approaches have demonstrated remarkable quality, their applicability is…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-14 Chaeyoung Jung , Suyeon Lee , Ji-Hoon Kim , Joon Son Chung

End-to-end automatic speech recognition systems represent the state of the art, but they rely on thousands of hours of manually annotated speech for training, as well as heavyweight computation for inference. Of course, this impedes…

Computation and Language · Computer Science 2022-11-22 Raphael Tang , Karun Kumar , Gefei Yang , Akshat Pandey , Yajie Mao , Vladislav Belyaev , Madhuri Emmadi , Craig Murray , Ferhan Ture , Jimmy Lin

All-neural end-to-end (E2E) automatic speech recognition (ASR) systems that use a single neural network to transduce audio to word sequences have been shown to achieve state-of-the-art results on several tasks. In this work, we examine the…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-28 Arun Narayanan , Rohit Prabhavalkar , Chung-Cheng Chiu , David Rybach , Tara N. Sainath , Trevor Strohman

Audio-visual speech enhancement system is regarded as one of promising solutions for isolating and enhancing speech of desired speaker. Typical methods focus on predicting clean speech spectrum via a naive convolution neural network based…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-01 Xinmeng Xu , Yang Wang , Jie Jia , Binbin Chen , Dejun Li

Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Although existing methods have achieved impressive performance,…

Multimedia · Computer Science 2026-03-03 Jiadong Wang , Ke Zhang , Xinyuan Qian , Ruijie Tao , Haizhou Li , Björn Schuller

Audio-visual speech enhancement system is regarded to be one of promising solutions for isolating and enhancing speech of desired speaker. Conventional methods focus on predicting clean speech spectrum via a naive convolution neural network…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-28 Xinmeng Xu , Jianjun Hao

This work presents our end-to-end (E2E) automatic speech recognition (ASR) model targetting at robust speech recognition, called Integraded speech Recognition with enhanced speech Input for Self-supervised learning representation (IRIS).…

Sound · Computer Science 2022-04-04 Xuankai Chang , Takashi Maekaku , Yuya Fujita , Shinji Watanabe

This paper describes an end-to-end (E2E) neural architecture for the audio rendering of small portions of display content on low resource personal computing devices. It is intended to address the problem of accessibility for vision-impaired…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-13 Liu Chen , Michael Deisher , Munir Georges

End-to-end (E2E) spoken language understanding (SLU) systems that generate a semantic parse from speech have become more promising recently. This approach uses a single model that utilizes audio and text representations from pre-trained…

Computation and Language · Computer Science 2023-07-25 Suyoun Kim , Akshat Shrivastava , Duc Le , Ju Lin , Ozlem Kalinli , Michael L. Seltzer

Due to the unprecedented breakthroughs brought about by deep learning, speech enhancement (SE) techniques have been developed rapidly and play an important role prior to acoustic modeling to mitigate noise effects on speech. To increase the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-15 Fu-An Chao , Shao-Wei Fan Jiang , Bi-Cheng Yan , Jeih-weih Hung , Berlin Chen

Speech enhancement can potentially benefit from the visual information from the target speaker, such as lip movement and facial expressions, because the visual aspect of speech is essentially unaffected by acoustic environment. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-24 Xinmeng Xu , Jianjun Hao

Speech enhancement algorithms based on deep learning have greatly surpassed their traditional counterparts and are now being considered for the task of removing acoustic echo from hands-free communication systems. This is a challenging…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-11 Jean-Marc Valin , Srikanth Tenneti , Karim Helwani , Umut Isik , Arvindh Krishnaswamy

Audio-Visual Target Speaker Extraction (AVTSE) aims to isolate a target speaker's voice in a multi-speaker environment with visual cues as auxiliary. Most of the existing AVTSE methods encode visual and audio features simultaneously,…

Sound · Computer Science 2025-11-13 Zixuan Li , Xueliang Zhang , Lei Miao , Zhipeng Yan , Ying Sun , Chong Zhu

Active speaker detection plays a vital role in human-machine interaction. Recently, a few end-to-end audiovisual frameworks emerged. However, these models' inference time was not explored and are not applicable for real-time applications…

Sound · Computer Science 2022-11-24 Fiseha B. Tesema , Zheyuan Lin , Shiqiang Zhu , Wei Song , Jason Gu , Hong Wu

Speaker diarization is well studied for constrained audios but little explored for challenging in-the-wild videos, which have more speakers, shorter utterances, and inconsistent on-screen speakers. We address this gap by proposing an…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-28 Zexu Pan , Gordon Wichern , François G. Germain , Aswin Subramanian , Jonathan Le Roux