English
Related papers

Related papers: Deep Ad-hoc Beamforming

200 papers

The spatial covariance matrix has been considered to be significant for beamformers. Standing upon the intersection of traditional beamformers and deep neural networks, we propose a causal neural beamformer paradigm called Embedding and…

Sound · Computer Science 2021-09-03 Andong Li , Wenzhe Liu , Chengshi Zheng , Xiaodong Li

This paper presents the details of the SRIB-LEAP submission to the ConferencingSpeech challenge 2021. The challenge involved the task of multi-channel speech enhancement to improve the quality of far field speech from microphone arrays in a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-25 R G Prithvi Raj , Rohit Kumar , M K Jayesh , Anurenjan Purushothaman , Sriram Ganapathy , M A Basha Shaik

Speech recognition in adverse real-world environments is highly affected by reverberation and nonstationary background noise. A well-known strategy to reduce such undesired signal components in multi-microphone scenarios is spatial…

Sound · Computer Science 2017-08-08 Hendrik Barfuss , Christian Huemmer , Andreas Schwarz , Walter Kellermann

Conventional sound source localization methods are mostly based on a single microphone array that consists of multiple microphones. They are usually formulated as the estimation of the direction of arrival problem. In this paper, we propose…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-18 Yijun Gong , Shupei Liu , Xiao-Lei Zhang

Although mask-based beamforming is a powerful speech enhancement approach, it often requires manual parameter tuning to handle moving speakers. Recently, this approach was augmented with an attention-based spatial covariance matrix…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-18 Marvin Tammen , Tsubasa Ochiai , Marc Delcroix , Tomohiro Nakatani , Shoko Araki , Simon Doclo

Enhancing the user's own-voice for head-worn microphone arrays is an important task in noisy environments to allow for easier speech communication and user-device interaction. However, a rarely addressed challenge is the change of the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-15 Wiebke Middelberg , Jung-Suk Lee , Saeed Bagheri Sereshki , Ali Aroudi , Vladimir Tourbabin , Daniel D. E. Wong

The performance of speech enhancement algorithms in a multi-speaker scenario depends on correctly identifying the target speaker to be enhanced. Auditory attention decoding (AAD) methods allow to identify the target speaker which the…

Sound · Computer Science 2020-05-12 Ali Aroudi , Marc Delcroix , Tomohiro Nakatani , Keisuke Kinoshita , Shoko Araki , Simon Doclo

The end-to-end approach for single-channel speech separation has been studied recently and shown promising results. This paper extended the previous approach and proposed a new end-to-end model for multi-channel speech separation. The…

Sound · Computer Science 2019-05-29 Rongzhi Gu , Jian Wu , Shi-Xiong Zhang , Lianwu Chen , Yong Xu , Meng Yu , Dan Su , Yuexian Zou , Dong Yu

Speech separation with several speakers is a challenging task because of the non-stationarity of the speech and the strong signal similarity between interferent sources. Current state-of-the-art solutions can separate well the different…

Signal Processing · Electrical Eng. & Systems 2021-02-09 Nicolas Furnon , Romain Serizel , Irina Illina , Slim Essid

Generating high-fidelity talking head video by fitting with the input audio sequence is a challenging problem that receives considerable attentions recently. In this paper, we address this problem with the aid of neural scene representation…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Yudong Guo , Keyu Chen , Sen Liang , Yong-Jin Liu , Hujun Bao , Juyong Zhang

Deep learning-based speech enhancement has seen huge improvements and recently also expanded to full band audio (48 kHz). However, many approaches have a rather high computational complexity and require big temporal buffers for real time…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-12 Hendrik Schröter , Alberto N. Escalante-B. , Tobias Rosenkranz , Andreas Maier

Acoustic echo cancellation (AEC) in full-duplex communication systems eliminates acoustic feedback. However, nonlinear distortions induced by audio devices, background noise, reverberation, and double-talk reduce the efficiency of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-29 Vinay Kothapally , Yong Xu , Meng Yu , Shi-Xiong Zhang , Dong Yu

The performance of traditional linear spatial filters for speech enhancement is constrained by the physical size and number of channels of microphone arrays. For instance, for large microphone distances and high frequencies, spatial…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-01 Alina Mannanova , Jakob Kienegger , Timo Gerkmann

While existing end-to-end beamformers achieve impressive performance in various front-end speech processing tasks, they usually encapsulate the whole process into a black box and thus lack adequate interpretability. As an attempt to fill…

Sound · Computer Science 2022-03-17 Andong Li , Guochen Yu , Chengshi Zheng , Xiaodong Li

The front-end module in multi-channel automatic speech recognition (ASR) systems mainly use microphone array techniques to produce enhanced signals in noisy conditions with reverberation and echos. Recently, neural network (NN) based…

Sound · Computer Science 2020-11-19 Yuxiang Kong , Jian Wu , Quandong Wang , Peng Gao , Weiji Zhuang , Yujun Wang , Lei Xie

The speech field is evolving to solve more challenging scenarios, such as multi-channel recordings with multiple simultaneous talkers. Given the many types of microphone setups out there, we present the UniX-Encoder. It's a universal…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-26 Zili Huang , Yiwen Shao , Shi-Xiong Zhang , Dong Yu

Recent studies have demonstrated that incorporating auxiliary information, such as speaker voiceprint or visual cues, can substantially improve Speech Enhancement (SE) performance. However, single-channel methods often yield suboptimal…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-06 Chihyun Liu , Jiaxuan Fan , Mingtung Sun , Michael Anthony , Mingsian R. Bai , Yu Tsao

Diffusion models have recently achieved impressive results in reconstructing images from noisy inputs, and similar ideas have been applied to speech enhancement by treating time-frequency representations as images. With the ubiquity of…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Renana Opochinsky , Sharon Gannot

Invariance to microphone array configuration is a rare attribute in neural beamformers. Filter-and-sum (FS) methods in this class define the target signal with respect to a reference channel. However, this not only complicates formulation…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-28 Anton Kovalyov , Kashyap Patel , Issa Panahi

Distant speech processing is a challenging task, especially when dealing with the cocktail party effect. Sound source separation is thus often required as a preprocessing step prior to speech recognition to improve the signal to distortion…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-06 Francois Grondin , Jean-Samuel Lauzon , Jonathan Vincent , Francois Michaud