English
Related papers

Related papers: Efficient Audio-Visual Speech Separation with Disc…

200 papers

Speech separation aims to separate individual voice from an audio mixture of multiple simultaneous talkers. Although audio-only approaches achieve satisfactory performance, they build on a strategy to handle the predefined conditions,…

Sound · Computer Science 2020-12-01 Peng Zhang , Jiaming Xu , Jing shi , Yunzhe Hao , Bo Xu

While existing Audio-Visual Speech Separation (AVSS) methods primarily concentrate on the audio-visual fusion strategy for two-speaker separation, they demonstrate a severe performance drop in the multi-speaker separation scenarios.…

Sound · Computer Science 2024-07-31 Tianrui Pan , Jie Liu , Bohan Wang , Jie Tang , Gangshan Wu

The goal of this paper is to learn strong lip reading models that can recognise speech in silent videos. Most prior works deal with the open-set visual speech recognition problem by adapting existing automatic speech recognition techniques…

Computer Vision and Pattern Recognition · Computer Science 2021-12-06 K R Prajwal , Triantafyllos Afouras , Andrew Zisserman

This paper presents Dolphin, a novel decoder-decoder architecture for energy-efficient processing of long contexts in language models. Our approach addresses the significant energy consumption and latency challenges inherent in on-device…

Computation and Language · Computer Science 2024-09-04 Wei Chen , Zhiyuan Li , Shuo Xin , Yihao Wang

We introduce AudioScopeV2, a state-of-the-art universal audio-visual on-screen sound separation system which is capable of learning to separate sounds and associate them with on-screen objects by looking at in-the-wild videos. We identify…

Sound · Computer Science 2022-07-22 Efthymios Tzinis , Scott Wisdom , Tal Remez , John R. Hershey

Audio-visual speech separation (AVSS) aims to extract a target speech signal from a mixed signal by leveraging both auditory and visual (lip movement) cues. However, most existing AVSS methods exhibit complex architectures and rely on…

Sound · Computer Science 2025-10-15 Wendi Sang , Kai Li , Runxuan Yang , Jianqiang Huang , Xiaolin Hu

Speech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources. Traditionally, these tasks have been…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-16 Daniel Michelsanti , Zheng-Hua Tan , Shi-Xiong Zhang , Yong Xu , Meng Yu , Dong Yu , Jesper Jensen

We present a joint audio-visual model for isolating a single speech signal from a mixture of sounds such as other speakers and background noise. Solving this task using only audio as input is extremely challenging and does not provide an…

This report introduces Dolphin, a large-scale multilingual automatic speech recognition (ASR) model that extends the Whisper architecture to support a wider range of languages. Our approach integrates in-house proprietary and open-source…

Computation and Language · Computer Science 2025-03-27 Yangyang Meng , Jinpeng Li , Guodong Lin , Yu Pu , Guanbo Wang , Hu Du , Zhiming Shao , Yukai Huang , Ke Li , Wei-Qiang Zhang

Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separation strategies prefer a straightforward fusion based on…

Sound · Computer Science 2022-03-08 Junwen Xiong , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha , Yanning Zhang

Audio-visual semantic segmentation (AVSS) represents an extension of the audio-visual segmentation (AVS) task, necessitating a semantic understanding of audio-visual scenes beyond merely identifying sound-emitting objects at the visual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yujian Lee , Peng Gao , Yongqi Xu , Wentao Fan

Visual speech recognition (VSR), also known as lip reading, is the task of recognizing speech from silent video. Despite significant advancements in VSR over recent decades, most existing methods pay limited attention to real-world visual…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Tianyue Wang , Shuang Yang , Shiguang Shan , Xilin Chen

Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large language models (Video-LLMs) can encounter many audio-centric…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Yuxin Guo , Shuailei Ma , Shijie Ma , Xiaoyi Bao , Chen-Wei Xie , Kecheng Zheng , Tingyu Weng , Siyang Sun , Yun Zheng , Wei Zou

In recent research, slight performance improvement is observed from automatic speech recognition systems to audio-visual speech recognition systems in the end-to-end framework with low-quality videos. Unmatching convergence rates and…

Computation and Language · Computer Science 2024-03-12 Yusheng Dai , Hang Chen , Jun Du , Xiaofei Ding , Ning Ding , Feijun Jiang , Chin-Hui Lee

Lip reading has witnessed unparalleled development in recent years thanks to deep learning and the availability of large-scale datasets. Despite the encouraging results achieved, the performance of lip reading, unfortunately, remains…

Computer Vision and Pattern Recognition · Computer Science 2019-11-27 Ya Zhao , Rui Xu , Xinchao Wang , Peng Hou , Haihong Tang , Mingli Song

Speech recognition is the technology that enables machines to interpret and process human speech, converting spoken language into text or commands. This technology is essential for applications such as virtual assistants, transcription…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-09 Xinyu Wang , Haotian Jiang , Haolin Huang , Yu Fang , Mengjie Xu , Qian Wang

Separating target speech from mixed signals containing flexible speaker quantities presents a challenging task. While existing methods demonstrate strong separation performance and noise robustness, they predominantly assume prior knowledge…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-18 Daning Zhang , Ying Wei

This paper focuses on designing a noise-robust end-to-end Audio-Visual Speech Recognition (AVSR) system. To this end, we propose Visual Context-driven Audio Feature Enhancement module (V-CAFE) to enhance the input noisy audio speech with a…

Sound · Computer Science 2022-07-14 Joanna Hong , Minsu Kim , Daehun Yoo , Yong Man Ro

Human auditory cortex excels at selectively suppressing background noise to focus on a target speaker. The process of selective attention in the brain is known to contextually exploit the available audio and visual cues to better focus on…

Sound · Computer Science 2018-09-12 Mandar Gogate , Ahsan Adeel , Ricard Marxer , Jon Barker , Amir Hussain

Accurate recognition of cocktail party speech containing overlapping speakers, noise and reverberation remains a highly challenging task to date. Motivated by the invariance of visual modality to acoustic signal corruption, an audio-visual…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-07 Guinan Li , Jiajun Deng , Mengzhe Geng , Zengrui Jin , Tianzi Wang , Shujie Hu , Mingyu Cui , Helen Meng , Xunying Liu
‹ Prev 1 2 3 10 Next ›