English
Related papers

Related papers: Speakerfilter-Pro: an improved target speaker extr…

200 papers

This paper proposes a new pitch estimator and a novel pitch tracker for speakers. We first decompose the sound signal into subbands using an auditory filterbank, assuming time-frequency sparsity of human speech. Instead of directly…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-03 Shoufeng Lin

The recent rapid development of auditory attention decoding (AAD) offers the possibility of using electroencephalography (EEG) as auxiliary information for target speaker extraction. However, effectively modeling long sequences of speech…

Sound · Computer Science 2025-01-06 Cunhang Fan , Youdian Gao , Zexu Pan , Jingjing Zhang , Hongyu Zhang , Jie Zhang , Zhao Lv

In this paper, we propose a new differentiable neural network alignment mechanism for text-dependent speaker verification which uses alignment models to produce a supervector representation of an utterance. Unlike previous works with…

Sound · Computer Science 2018-12-27 Victoria Mingote , Antonio Miguel , Alfonso Ortega , Eduardo Lleida

High quality speech capture has been widely studied for both voice communication and human computer interface reasons. To improve the capture performance, we can often find multi-microphone speech enhancement techniques deployed on various…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-15 Yang Yang , Shao-Fu Shih , Hakan Erdogan , Jamie Menjay Lin , Chehung Lee , Yunpeng Li , George Sung , Matthias Grundmann

Target speaker extraction (TSE) is a technique for isolating a target speaker's voice from mixed speech using auxiliary features associated with the target speaker. It is another attempt at addressing the cocktail party problem and is…

Sound · Computer Science 2024-11-26 Chang Sun , Bo Qin

We address monaural multi-speaker-image separation in reverberant conditions, aiming at separating mixed speakers but preserving the reverberation of each speaker. A straightforward approach for this task is to directly train end-to-end DNN…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-08 Jingqi Sun , Shulin He , Ruizhe Pang , Zhong-Qiu Wang

We present a novel source separation model to decompose asingle-channel speech signal into two speech segments belonging to two different speakers. The proposed model is a neural network based on residual blocks, and uses learnt speaker…

Sound · Computer Science 2019-06-25 Shuo Liu , Gil Keren , Björn Schuller

Real-time target speaker extraction (TSE) is intended to extract the desired speaker's voice from the observed mixture of multiple speakers in a streaming manner. Implementing real-time TSE is challenging as the computational complexity…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-03 Hiroshi Sato , Takafumi Moriya , Masato Mimura , Shota Horiguchi , Tsubasa Ochiai , Takanori Ashihara , Atsushi Ando , Kentaro Shinayama , Marc Delcroix

The dual-path RNN (DPRNN) was proposed to more effectively model extremely long sequences for speech separation in the time domain. By splitting long sequences to smaller chunks and applying intra-chunk and inter-chunk RNNs, the DPRNN…

Sound · Computer Science 2021-07-13 Xiaohuai Le , Hongsheng Chen , Kai Chen , Jing Lu

Recent research in speaker verification has increasingly focused on achieving robust and reliable recognition under challenging channel conditions and noisy environments. Identifying speakers in radio communications is particularly…

Sound · Computer Science 2024-06-18 Wenhao Yang , Jianguo Wei , Wenhuan Lu , Lei Li , Xugang Lu

Target speaker extraction (TSE) aims to isolate a specific speaker's speech from a mixture using speaker enrollment as a reference. While most existing approaches are discriminative, recent generative methods for TSE achieve strong results.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-21 Aviv Navon , Aviv Shamsian , Yael Segal-Feldman , Neta Glazer , Gil Hetz , Joseph Keshet

A three-stage approach is proposed for speaker counting and speech separation in noisy and reverberant environments. In the spatial feature extraction, a spatial coherence matrix (SCM) is computed using whitened relative transfer functions…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-08 Yicheng Hsu , Mingsian Bai

We propose a novel Neural Steering technique that adapts the target area of a spatial-aware multi-microphone sound source separation algorithm during inference without the necessity of retraining the deep neural network (DNN). To achieve…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-23 Martin Strauss , Wolfgang Mack , María Luis Valero , Okan Köpüklü

Large-scale pre-trained self-supervised learning (SSL) models have shown remarkable advancements in speech-related tasks. However, the utilization of these models in complex multi-talker scenarios, such as extracting a target speaker in a…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-21 Junyi Peng , Marc Delcroix , Tsubasa Ochiai , Oldrich Plchot , Takanori Ashihara , Shoko Araki , Jan Cernocky

Target speech extraction is a technique to extract the target speaker's voice from mixture signals using a pre-recorded enrollment utterance that characterize the voice characteristics of the target speaker. One major difficulty of target…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-17 Hiroshi Sato , Tsubasa Ochiai , Marc Delcroix , Keisuke Kinoshita , Takafumi Moriya , Naoki Makishima , Mana Ihori , Tomohiro Tanaka , Ryo Masumura

Spectral subtraction, widely used for its simplicity, has been employed to address the Robot Ego Speech Filtering (RESF) problem for detecting speech contents of human interruption from robot's single-channel microphone recordings when it…

Robotics · Computer Science 2024-09-11 Yue Li , Koen V. Hindriks , Florian A. Kunneman

In this study, we propose the global context guided channel and time-frequency transformations to model the long-range, non-local time-frequency dependencies and channel variances in speaker representations. We use the global context…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-10 Wei Xia , John H. L. Hansen

The task of estimating the maximum number of concurrent speakers from single channel mixtures is important for various audio-based applications, such as blind source separation, speaker diarisation, audio surveillance or auditory scene…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-05 Fabian-Robert Stöter , Soumitro Chakrabarty , Bernd Edler , Emanuël A. P. Habets

Recently, a spatially selective non-linear filter (SSF) has been proposed for target speaker extraction, using the target direction-of-arrival (DOA) as a spatial cue. Since learned intermediate features are tied to the microphone geometry,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-19 Jiatong Li , Wiebke Middelberg , Simon Doclo

Multi-channel speech separation in dynamic environments is challenging as time-varying spatial and spectral features evolve at different temporal scales. Existing methods typically employ sequential architectures, forcing a single network…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-27 Yuzhu Wang , Archontis Politis , Konstantinos Drossos , Tuomas Virtanen