中文
相关论文

相关论文: Interpretable Binaural Deep Beamforming Guided by …

200 篇论文

This work addresses the problem of block-online processing for multi-channel speech enhancement. Such processing is vital in scenarios with moving speakers and/or when very short utterances are processed, e.g., in voice assistant scenarios.…

声音 · 计算机科学 2020-05-27 Jiri Malek , Zbynek Koldovsky , Marek Bohac

We propose BeamTransformer, an efficient architecture to leverage beamformer's edge in spatial filtering and transformer's capability in context sequence modeling. BeamTransformer seeks to optimize modeling of sequential relationship among…

声音 · 计算机科学 2021-09-10 Siqi Zheng , Shiliang Zhang , Weilong Huang , Qian Chen , Hongbin Suo , Ming Lei , Jinwei Feng , Zhijie Yan

This paper presents a novel machine-hearing system that exploits deep neural networks (DNNs) and head movements for robust binaural localisation of multiple sources in reverberant environments. DNNs are used to learn the relationship…

音频与语音处理 · 电气工程与系统科学 2019-04-08 Ning Ma , Tobias May , Guy J. Brown

This study presents a deep-learning framework for controlling multichannel acoustic feedback in audio devices. Traditional digital signal processing methods struggle with convergence when dealing with highly correlated noise such as…

声音 · 计算机科学 2025-05-30 Yuan-Kuei Wu , Juan Azcarreta , Kashyap Patel , Buye Xu , Jung-Suk Lee , Sanha Lee , Ashutosh Pandey

The learning of interpretable representations from raw data presents significant challenges for time series data like speech. In this work, we propose a relevance weighting scheme that allows the interpretation of the speech representations…

音频与语音处理 · 电气工程与系统科学 2020-11-05 Purvi Agrawal , Sriram Ganapathy

Humans can robustly recognize and localize objects by integrating visual and auditory cues. While machines are able to do the same now with images, less work has been done with sounds. This work develops an approach for dense semantic…

计算机视觉与模式识别 · 计算机科学 2020-03-10 Arun Balajee Vasudevan , Dengxin Dai , Luc Van Gool

Beamforming techniques are utilized in millimeter wave (mmWave) communication to address the inherent path loss limitation, thereby establishing and maintaining reliable connections. However, adopting standard defined beamforming approach…

网络与互联网体系结构 · 计算机科学 2025-09-16 Muhammad Baqer Mollah , Honggang Wang , Hua Fang

In many multi-microphone algorithms, an estimate of the relative transfer functions (RTFs) of the desired speaker is required. Recently, a computationally efficient RTF vector estimation method was proposed for acoustic sensor networks,…

音频与语音处理 · 电气工程与系统科学 2023-04-07 Wiebke Middelberg , Simon Doclo

Humans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene…

声音 · 计算机科学 2022-03-01 Dengxin Dai , Arun Balajee Vasudevan , Jiri Matas , Luc Van Gool

Speaker localization for binaural microphone arrays has been widely studied for applications such as speech communication, video conferencing, and robot audition. Many methods developed for this task, including the direct path dominance…

音频与语音处理 · 电气工程与系统科学 2023-11-01 Yanir Maymon , Israel Nelken , Boaz Rafaely

This paper investigates deep learning techniques to predict transmit beamforming based on only historical channel data without current channel information in the multiuser multiple-input-single-output downlink. This will significantly…

信息论 · 计算机科学 2023-02-03 Juping Zhang , Gan Zheng , Yangyishi Zhang , Ioannis Krikidis , Kai-Kit Wong

With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical. Existing methods for audio-visual deepfake detection typically rely on cross-modal inconsistencies. In singing, rhythmic…

人工智能 · 计算机科学 2026-05-28 Ke Liu , Jiwei Wei , Wenyu Zhang , Shuchang Zhou , Ruikun Chai , Yutao Dai , Chaoning Zhang , Yang Yang

Most existing beamforming methods are frequency-domain methods, and are designed for enhancing a farfield target source over a narrow frequency band. They have found diverse applications and are still under active development. However, they…

音频与语音处理 · 电气工程与系统科学 2021-05-19 Fei Ma , Thushara D. Abhayapala , Prasanga N. Samarasinghe

In this paper, we introduce a spectral-domain inverse filtering approach for single-channel speech de-reverberation using deep convolutional neural network (CNN). The main goal is to better handle realistic reverberant conditions where the…

声音 · 计算机科学 2020-10-16 Hanwook Chung , Vikrant Singh Tomar , Benoit Champagne

Head Related Transfer Functions (HRTFs) play a crucial role in creating immersive spatial audio experiences. However, HRTFs differ significantly from person to person, and traditional methods for estimating personalized HRTFs are expensive,…

音频与语音处理 · 电气工程与系统科学 2023-11-08 Vivek Jayaram , Ira Kemelmacher-Shlizerman , Steven M. Seitz

This paper presents two single channel speech dereverberation methods to enhance the quality of speech signals that have been recorded in an enclosed space. For both methods, the room acoustics are modeled using a nonnegative approximation…

声音 · 计算机科学 2017-09-19 Nasser Mohammadiha , Simon Doclo

Multi-channel acoustic signal processing is a well-established and powerful tool to exploit the spatial diversity between a target signal and non-target or noise sources for signal enhancement. However, the textbook solutions for optimal…

音频与语音处理 · 电气工程与系统科学 2025-01-14 Reinhold Haeb-Umbach , Tomohiro Nakatani , Marc Delcroix , Christoph Boeddeker , Tsubasa Ochiai

A target recognition framework relying on near-field integrated sensing and communication (ISAC) systems is proposed. By exploiting the distance-dependent spatial signatures provided by the near-field spherical wavefront, high-accuracy…

信号处理 · 电气工程与系统科学 2026-03-17 Zongyao Zhao , Zhaolin Wang , Lincong Han , Jing Jin , Kaibin Huang

The performance of deep learning-based multi-channel speech enhancement methods often deteriorates when the geometric parameters of the microphone array change. Traditional approaches to mitigate this issue typically involve training on…

音频与语音处理 · 电气工程与系统科学 2025-04-03 Tianqin Zheng , Jilu Jin , Hanchen Pei , Gongping Huang , Jingdong Chen , Jacob Benesty

Speaker Diarization (SD) aims at grouping speech segments that belong to the same speaker. This task is required in many speech-processing applications, such as rich meeting transcription. In this context, distant microphone arrays usually…

声音 · 计算机科学 2024-06-06 Theo Mariotte , Anthony Larcher , Silvio Montresor , Jean-Hugh Thomas