English
Related papers

Related papers: Flexible Multi-Channel Target Speaker Extraction U…

200 papers

Recently, speech enhancement technologies that are based on deep learning have received considerable research attention. If the spatial information in microphone signals is exploited, microphone arrays can be advantageous under some adverse…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-19 Yicheng Hsu , Yonghan Lee , Mingsian R. Bai

Target speech separation refers to extracting the target speaker's speech from mixed signals. Despite the recent advances in deep learning based close-talk speech separation, the applications to real-world are still an open issue. Two main…

Sound · Computer Science 2020-01-03 Rongzhi Gu , Yuexian Zou

Multi-channel target speaker extraction (MC-TSE) aims to extract a target speaker's voice from multi-speaker signals captured by multiple microphones. Existing methods often rely on auxiliary clues such as direction-of-arrival (DOA) or…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-20 Tongtao Ling , Shulin He , Pengjie Shen , Zhong-Qiu Wang

In a recent work on direction-of-arrival (DOA) estimation of multiple speakers with convolutional neural networks (CNNs), the phase component of short-time Fourier transform (STFT) coefficients of the microphone signal is given as input and…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-18 Soumitro Chakrabarty , Emanuël A. P. Habets

This paper presents a robust multi-channel speaker extraction algorithm designed to handle inaccuracies in reference information. While existing approaches often rely solely on either spatial or spectral cues to identify the target speaker,…

Sound · Computer Science 2025-12-24 Aviad Eisenberg , Sharon Gannot , Shlomo E. Chazan

Target Speaker Extraction (TSE) plays a critical role in enhancing speech signals in noisy and multi-speaker environments. This paper presents an end-to-end TSE model that incorporates Direction of Arrival (DOA) and beamwidth embeddings to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-29 Kangqi Jing , Wenbin Zhang , Yu Gao

Multi-channel speech enhancement extracts speech using multiple microphones that capture spatial cues. Effectively utilizing directional information is key for multi-channel enhancement. Deep learning shows great potential on multi-channel…

Sound · Computer Science 2023-09-21 Jiahui Pan , Pengjie Shen , Hui Zhang , Xueliang Zhang

Spatial filters can exploit deep-learning-based speech enhancement models to increase their reliability in scenarios with multiple speech sources scenarios. To further improve speech quality, it is common to perform postfiltering on the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-13 François Grondin , Caleb Rascón

Recent works on deep non-linear spatially selective filters demonstrate exceptional enhancement performance with computationally lightweight architectures for stationary speakers of known directions. However, to maintain this performance in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-08 Jakob Kienegger , Alina Mannanova , Huajian Fang , Timo Gerkmann

Recently, a relative transfer function (RTF)-vector-based method has been proposed to estimate the direction of arrival (DOA) of a target speaker for a binaural hearing aid setup, assuming the availability of external microphones. This…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-11 Daniel Fejgin , Simon Doclo

Multi-channel speech enhancement utilizes spatial information from multiple microphones to extract the target speech. However, most existing methods do not explicitly model spatial cues, instead relying on implicit learning from…

Sound · Computer Science 2023-09-20 Jiahui Pan , Shulin He , Hui Zhang , Xueliang Zhang

Parametric sound field synthesis methods, such as the Spatial Decomposition Method (SDM) and Higher-Order Spatial Impulse Response Rendering (HO-SIRR), are widely used for the analysis and auralization of sound fields. This paper studies…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-04 Alan Pawlak , Hyunkook Lee , Aki Mäkivirta , Thomas Lund

In this paper, we attempt to study the conditioning of the Spherical Harmonic Matrix (SHM), which is widely used in the discrete, limited order orthogonal representation of sound fields. SHM's has been widely used in the audio applications…

Audio and Speech Processing · Electrical Eng. & Systems 2018-03-07 C Sandeep Reddy , Rajesh M Hegde

Direction-of-Arrival (DOA) estimation in sensor arrays faces limitations under demanding conditions, including low signal-to-noise ratio, single-snapshot scenarios, coherent sources, and unknown source counts. Conventional beamforming…

Signal Processing · Electrical Eng. & Systems 2025-10-14 Xuyao Deng , Yong Dou , Kele Xu

Speaker diarization is the task of answering Who spoke and when? in an audio stream. Pipeline systems rely on speech segmentation to extract speakers' segments and achieve robust speaker diarization. This paper proposes a common framework…

Sound · Computer Science 2023-06-08 Théo Mariotte , Anthony Larcher , Silvio Montrésor , Jean-Hugh Thomas

We propose a speaker selection mechanism (SSM) for the training of an end-to-end beamforming neural network, based on recent findings that a listener usually looks to the target speaker with a certain undershot angle. The mechanism allows…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-25 Luan Vinícius Fiorio , Bruno Defraene , Johan David , Alex Young , Frans Widdershoven , Wim van Houtum , Ronald M. Aarts

In this paper, we propose an effective and robust method of spatial feature extraction for acoustic scene analysis utilizing partially synchronized and/or closely located distributed microphones. In the proposed method, a new cepstrum…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-22 Keisuke Imoto

Spotforming is a target-speaker extraction technique that uses multiple microphone arrays. This method applies beamforming (BF) to each microphone array, and the common components among the BF outputs are estimated as the target source.…

Sound · Computer Science 2024-07-15 Shoma Ayano , Li Li , Shogo Seki , Daichi Kitamura

We consider the task of region-based source separation of reverberant multi-microphone recordings. We assume pre-defined spatial regions with a single active source per region. The objective is to estimate the signals from the individual…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-14 Julian Wechsler , Srikanth Raj Chetupalli , Wolfgang Mack , Emanuël A. P. Habets

In hearing aid applications, an important objective is to accurately estimate the direction of arrival (DOA) of multiple speakers in noisy and reverberant environments. Recently, we proposed a binaural DOA estimation method, where the DOAs…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-11 Daniel Fejgin , Simon Doclo