中文
相关论文

相关论文: RAVSS: Robust Audio-Visual Speech Separation in Mu…

200 篇论文

Active speaker detection (ASD) systems are important modules for analyzing multi-talker conversations. They aim to detect which speakers or none are talking in a visual scene at any given time. Existing research on ASD does not agree on the…

声音 · 计算机科学 2022-07-12 Abudukelimu Wuerkaixi , You Zhang , Zhiyao Duan , Changshui Zhang

Numerous studies have investigated the effectiveness of audio-visual multimodal learning for speech enhancement (AVSE) tasks, seeking a solution that uses visual data as auxiliary and complementary input to reduce the noise of noisy speech…

音频与语音处理 · 电气工程与系统科学 2022-02-02 Shang-Yi Chuang , Hsin-Min Wang , Yu Tsao

AV-HuBERT, a multi-modal self-supervised learning model, has been shown to be effective for categorical problems such as automatic speech recognition and lip-reading. This suggests that useful audio-visual speech representations can be…

音频与语音处理 · 电气工程与系统科学 2023-06-02 I-Chun Chern , Kuo-Hsuan Hung , Yi-Ting Chen , Tassadaq Hussain , Mandar Gogate , Amir Hussain , Yu Tsao , Jen-Cheng Hou

Speech enhancement systems are typically trained using pairs of clean and noisy speech. In audio-visual speech enhancement (AVSE), there is not as much ground-truth clean data available; most audio-visual datasets are collected in…

音频与语音处理 · 电气工程与系统科学 2024-11-05 Ju-Chieh Chou , Chung-Ming Chien , Karen Livescu

This work presents an extensive and detailed study on Audio-Visual Speech Recognition (AVSR) for five widely spoken languages: Chinese, Spanish, English, Arabic, and French. We have collected large-scale datasets for each language except…

计算与语言 · 计算机科学 2024-06-04 Sanath Narayan , Yasser Abdelaziz Dahou Djilali , Ankit Singh , Eustache Le Bihan , Hakim Hacid

Automatic speaker verification (ASV) is the process to recognize persons using voice as biometric. The ASV systems show considerable recognition performance with sufficient amount of speech from matched condition. One of the crucial…

多媒体 · 计算机科学 2018-12-04 Arnab Poddar , Md Sahidullah , Goutam Saha

The practical deployment of Audio-Visual Speech Recognition (AVSR) systems is fundamentally challenged by significant performance degradation in real-world environments, characterized by unpredictable acoustic noise and visual interference.…

音频与语音处理 · 电气工程与系统科学 2025-12-17 Sungnyun Kim

Automatic speech recognition (ASR) of overlapped speech remains a highly challenging task to date. To this end, multi-channel microphone array data are widely used in state-of-the-art ASR systems. Motivated by the invariance of visual…

音频与语音处理 · 电气工程与系统科学 2020-11-19 Jianwei Yu , Bo Wu , Rongzhi Gu , Shi-Xiong Zhang , Lianwu Chen , Yong Xu. Meng Yu , Dan Su , Dong Yu , Xunying Liu , Helen Meng

We develop an end-to-end system for multi-channel, multi-speaker automatic speech recognition. We propose a frontend for joint source separation and dereverberation based on the independent vector analysis (IVA) paradigm. It uses the fast…

音频与语音处理 · 电气工程与系统科学 2022-04-04 Robin Scheibler , Wangyou Zhang , Xuankai Chang , Shinji Watanabe , Yanmin Qian

Audio-visual speech recognition (AVSR) incorporates auditory and visual modalities to improve recognition accuracy, particularly in noisy environments where audio-only speech systems are insufficient. While previous research has largely…

音频与语音处理 · 电气工程与系统科学 2025-05-01 Sungnyun Kim , Sungwoo Cho , Sangmin Bae , Kangwook Jang , Se-Young Yun

We present a frontend for improving robustness of automatic speech recognition (ASR), that jointly implements three modules within a single model: acoustic echo cancellation, speech enhancement, and speech separation. This is achieved by…

音频与语音处理 · 电气工程与系统科学 2021-11-22 Tom O'Malley , Arun Narayanan , Quan Wang , Alex Park , James Walker , Nathan Howard

Active speaker detection is an important component in video analysis algorithms for applications such as speaker diarization, video re-targeting for meetings, speech enhancement, and human-robot interaction. The absence of a large,…

With the advance in self-supervised learning for audio and visual modalities, it has become possible to learn a robust audio-visual speech representation. This would be beneficial for improving the audio-visual speech recognition (AVSR)…

图像与视频处理 · 电气工程与系统科学 2022-07-12 Zi-Qiang Zhang , Jie Zhang , Jian-Shu Zhang , Ming-Hui Wu , Xin Fang , Li-Rong Dai

Speaker diarization consists of assigning speech signals to people engaged in a dialogue. An audio-visual spatiotemporal diarization model is proposed. The model is well suited for challenging scenarios that consist of several participants…

计算机视觉与模式识别 · 计算机科学 2018-10-15 Israel D. Gebru , Silèye Ba , Xiaofei Li , Radu Horaud

Real-time single-channel speech separation aims to unmix an audio stream captured from a single microphone that contains multiple people talking at once, environmental noise, and reverberation into multiple de-reverberated and noise-free…

音频与语音处理 · 电气工程与系统科学 2023-04-18 Julian Neri , Sebastian Braun

Recent years have witnessed the extraordinary development of automatic speaker verification (ASV). However, previous works show that state-of-the-art ASV models are seriously vulnerable to voice spoofing attacks, and the recently proposed…

声音 · 计算机科学 2022-06-22 Haibin Wu , Jiawen Kang , Lingwei Meng , Yang Zhang , Xixin Wu , Zhiyong Wu , Hung-yi Lee , Helen Meng

The goal of this paper is speech separation and enhancement in multi-speaker and noisy environments using a combination of different modalities. Previous works have shown good performance when conditioning on temporal or static visual…

音频与语音处理 · 电气工程与系统科学 2025-01-06 Akam Rahimi , Triantafyllos Afouras , Andrew Zisserman

This paper studies audio-visual noise suppression for egocentric videos -- where the speaker is not captured in the video. Instead, potential noise sources are visible on screen with the camera emulating the off-screen speaker's view of the…

声音 · 计算机科学 2023-05-04 Roshan Sharma , Weipeng He , Ju Lin , Egor Lakomkin , Yang Liu , Kaustubh Kalgaonkar

Speech separation has been successfully applied as a frontend processing module of conversation transcription systems thanks to its ability to handle overlapped speech and its flexibility to combine with downstream tasks such as automatic…

音频与语音处理 · 电气工程与系统科学 2021-07-06 Jian Wu , Zhuo Chen , Sanyuan Chen , Yu Wu , Takuya Yoshioka , Naoyuki Kanda , Shujie Liu , Jinyu Li

The ability to capture and segment sounding objects in dynamic visual scenes is crucial for the development of Audio-Visual Segmentation (AVS) tasks. While significant progress has been made in this area, the interaction between audio and…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Kai Peng , Yunzhe Shen , Miao Zhang , Leiye Liu , Yidong Han , Wei Ji , Jingjing Li , Yongri Piao , Huchuan Lu