中文
相关论文

相关论文: Robust Audio-Visual Target Speaker Extraction with…

200 篇论文

Audio-visual feature synchronization for real-time speech enhancement in hearing aids represents a progressive approach to improving speech intelligibility and user experience, particularly in strong noisy backgrounds. This approach…

音频与语音处理 · 电气工程与系统科学 2025-08-28 Nasir Saleem , Mandar Gogate , Kia Dashtipour , Adeel Hussain , Usman Anwar , Adewale Adetomi , Tughrul Arslan , Amir Hussain

Separating target speech from mixed signals containing flexible speaker quantities presents a challenging task. While existing methods demonstrate strong separation performance and noise robustness, they predominantly assume prior knowledge…

音频与语音处理 · 电气工程与系统科学 2025-07-18 Daning Zhang , Ying Wei

Audio-Visual Target Speaker Extraction (AV-TSE) aims to mimic the human ability to enhance auditory perception using visual cues. Although numerous models have been proposed recently, most of them estimate target signals by primarily…

声音 · 计算机科学 2025-04-02 Wenxuan Wu , Xueyuan Chen , Shuai Wang , Jiadong Wang , Lingwei Meng , Xixin Wu , Helen Meng , Haizhou Li

Real-time audio-visual speech enhancement (AVSE) is a key enabler for immersive and interactive multimedia services, yet its performance is tightly constrained by network latency, uplink capacity, and computational delay. This paper…

声音 · 计算机科学 2026-04-29 Anis Hamadouche , Haifeng Luo , Mathini Sellathurai , Amir Hussain , Tharm Ratnarajah

Previous studies have confirmed the effectiveness of incorporating visual information into speech enhancement (SE) systems. Despite improved denoising performance, two problems may be encountered when implementing an audio-visual SE (AVSE)…

音频与语音处理 · 电气工程与系统科学 2020-08-19 Shang-Yi Chuang , Yu Tsao , Chen-Chou Lo , Hsin-Min Wang

Automatic speech recognition can potentially benefit from the lip motion patterns, complementing acoustic speech to improve the overall recognition performance, particularly in noise. In this paper we propose an audio-visual fusion strategy…

音频与语音处理 · 电气工程与系统科学 2019-05-02 George Sterpu , Christian Saam , Naomi Harte

Audio-visual speech enhancement (AV-SE) aims to enhance degraded speech along with extra visual information such as lip videos, and has been shown to be more effective than audio-only speech enhancement. This paper proposes the…

音频与语音处理 · 电气工程与系统科学 2023-11-21 Rui-Chen Zheng , Yang Ai , Zhen-Hua Ling

With the growing success of multi-modal learning, research on the robustness of multi-modal models, especially when facing situations with missing modalities, is receiving increased attention. Nevertheless, previous studies in this domain…

人工智能 · 计算机科学 2023-10-11 Siting Li , Chenzhuang Du , Yue Zhao , Yu Huang , Hang Zhao

It is now well established from a variety of studies that there is a significant benefit from combining video and audio data in detecting active speakers. However, either of the modalities can potentially mislead audiovisual fusion by…

The primary goal of multi-modal TSE (MTSE) is to extract a target speaker from a speech mixture using complementary information from different modalities, such as audio enrolment and visual feeds corresponding to the target speaker. MTSE…

音频与语音处理 · 电气工程与系统科学 2025-07-10 Srikanth Korse , Mohamed Elminshawi , Emanuel A. P. Habets , Srikanth Raj Chetupalli

Speaker verification has been widely explored using speech signals, which has shown significant improvement using deep models. Recently, there has been a surge in exploring faces and voices as they can offer more complementary and…

声音 · 计算机科学 2023-09-29 R. Gnana Praveen , Jahangir Alam

Audio-visual speech recognition (AVSR) typically improves recognition accuracy in noisy environments by integrating noise-immune visual cues with audio signals. Nevertheless, high-noise audio inputs are prone to introducing adverse…

音频与语音处理 · 电气工程与系统科学 2026-03-09 Linzhi Wu , Xingyu Zhang , Hao Yuan , Yakun Zhang , Changyan Zheng , Liang Xie , Tiejun Liu , Erwei Yin

Utilizing the sensor characteristics of the audio, visible camera, and thermal camera, the robustness of person recognition can be enhanced. Existing multimodal person recognition frameworks are primarily formulated assuming that multimodal…

多媒体 · 计算机科学 2022-10-25 Vijay John , Yasutomo Kawanishi

Automatic emotion recognition (ER) has recently gained lot of interest due to its potential in many real-world applications. In this context, multimodal approaches have been shown to improve performance (over unimodal approaches) by…

计算机视觉与模式识别 · 计算机科学 2022-09-20 R Gnana Praveen , Eric Granger , Patrick Cardinal

Previous Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advanced back-end recognition systems often hit performance…

Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Although existing methods have achieved impressive performance,…

多媒体 · 计算机科学 2026-03-03 Jiadong Wang , Ke Zhang , Xinyuan Qian , Ruijie Tao , Haizhou Li , Björn Schuller

Audiovisual active speaker detection (ASD) in egocentric recordings is challenged by frequent occlusions, motion blur, and audio interference, which undermine the discernability of temporal synchrony between lip movement and speech.…

多媒体 · 计算机科学 2025-08-15 Jason Clarke , Yoshihiko Gotoh , Stefan Goetze

This paper describes our audio-quality-based multi-strategy approach for the audio-visual target speaker extraction (AVTSE) task in the Multi-modal Information based Speech Processing (MISP) 2023 Challenge. Specifically, our approach adopts…

声音 · 计算机科学 2024-03-08 Runduo Han , Xiaopeng Yan , Weiming Xu , Pengcheng Guo , Jiayao Sun , He Wang , Quan Lu , Ning Jiang , Lei Xie

Audio-visual speech enhancement (AVSE) has been found to be particularly useful at low signal-to-noise (SNR) ratios due to the immunity of the visual features to acoustic noise. However, a significant gap exists in AVSE methods tailored to…

音频与语音处理 · 电气工程与系统科学 2025-10-21 Danielle Yaffe , Ferdinand Campe , Prachi Sharma , Dorothea Kolossa , Boaz Rafaely

Audio-visual target speaker extraction (AV-TSE) aims to extract the specific person's speech from the audio mixture given auxiliary visual cues. Previous methods usually search for the target voice through speech-lip synchronization.…

音频与语音处理 · 电气工程与系统科学 2025-03-04 Ruijie Tao , Xinyuan Qian , Yidi Jiang , Junjie Li , Jiadong Wang , Haizhou Li