English
Related papers

Related papers: Flexible Multi-Channel Target Speaker Extraction U…

200 papers

In this study, we propose the global context guided channel and time-frequency transformations to model the long-range, non-local time-frequency dependencies and channel variances in speaker representations. We use the global context…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-10 Wei Xia , John H. L. Hansen

We propose a separation guided speaker diarization (SGSD) approach by fully utilizing a complementarity of speech separation and speaker clustering. Since the conventional clustering-based speaker diarization (CSD) approach cannot well…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-07 Shu-Tong Niu , Jun Du , Lei Sun , Chin-Hui Lee

A novel unified frequency diverse array (FDA) and spatial data focusing (SDF) approach is proposed to simultaneously overcome time-variance and precision constraints of conventional FDA in geocasting, i.e., spatially confined broadcasting,…

Signal Processing · Electrical Eng. & Systems 2023-02-27 Guylian Molineaux , Francois Horlin , Philippe De Doncker , Julien Sarrazin

Direction-of-arrival estimation of multiple speakers in a room is an important task for a wide range of applications. In particular, challenging environments with moving speakers, reverberation and noise, lead to significant performance…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-24 Daniel A. Mitchell , Boaz Rafaely , Anurag Kumar , Vladimir Tourbabin

In this paper, we propose an effective and robust method for acoustic scene analysis based on spatial information extracted from partially synchronized and/or closely located distributed microphones. In the proposed method, to extract…

Sound · Computer Science 2018-07-10 Keisuke Imoto

Neural source-filter (NSF) models are deep neural networks that produce waveforms given input acoustic features. They use dilated-convolution-based neural filter modules to filter sine-based excitation for waveform generation, which is…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-28 Xin Wang , Junichi Yamagishi

Target speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current end-to-end neural networks require large model parameters…

Sound · Computer Science 2024-01-08 Shulin He , Jinjiang liu , Hao Li , Yang Yang , Fei Chen , Xueliang Zhang

We propose a novel Neural Steering technique that adapts the target area of a spatial-aware multi-microphone sound source separation algorithm during inference without the necessity of retraining the deep neural network (DNN). To achieve…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-23 Martin Strauss , Wolfgang Mack , María Luis Valero , Okan Köpüklü

This paper proposes a model that integrates sub-band processing and deep filtering to fully exploit information from the target time-frequency (TF) bin and its surrounding TF bins for single-channel speech enhancement. The sub-band module…

Sound · Computer Science 2025-06-03 Shenghui Lu , Hukai Huang , Jinanglong Yao , Kaidi Wang , Qingyang Hong , Lin Li

In this paper, we introduce spatial attention for refining the information in multi-direction neural beamformer for far-field automatic speech recognition. Previous approaches of neural beamformers with multiple look directions, such as the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-10 Weipeng He , Lu Lu , Biqiao Zhang , Jay Mahadeokar , Kaustubh Kalgaonkar , Christian Fuegen

We propose a multichannel-to-multichannel target sound extraction (M2M-TSE) framework for separating multichannel target signals from a multichannel mixture of sound sources. Target sound extraction (TSE) isolates a specific target signal…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-20 Dayun Choi , Jung-Woo Choi

Deep Learning models have become potential candidates for auditory neuroscience research, thanks to their recent successes on a variety of auditory tasks. Yet, these models often lack interpretability to fully understand the exact…

Sound · Computer Science 2021-08-04 Rachid Riad , Julien Karadayi , Anne-Catherine Bachoud-Lévi , Emmanuel Dupoux

Spherical microphone arrays (SMAs) and spherical loudspeaker arrays (SLAs) facilitate the study of room acoustics due to the three-dimensional analysis they provide. More recently, systems that combine both arrays, referred to as…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-09 Hai Morgenstern , Boaz Rafaely , Markus Noisternig

Previous studies have confirmed the adverse impact of fading correlation on the mutual information (MI) of two-dimensional (2D) multiple-input multiple-output (MIMO) systems. More recently, the trend is to enhance the system performance by…

Information Theory · Computer Science 2016-04-21 Qurrat-Ul-Ain Nadeem , Abla Kammoun , Merouane Debbah , Mohamed-Slim Alouini

This paper introduces a novel methodology leveraging differentiable programming to design efficient, constrained adaptive non-uniform Linear Differential Microphone Arrays (LDMAs) with reduced implementation costs. Utilizing an automatic…

Sound · Computer Science 2024-12-09 Siminfar Samakoush Galougah , Ramani Duraiswami

Traditional speech separation and speaker diarization approaches rely on prior knowledge of target speakers or a predetermined number of participants in audio signals. To address these limitations, recent advances focus on developing…

Speech foundation models (SFMs) have demonstrated strong performance across a variety of downstream tasks, including speech intelligibility prediction for hearing-impaired people (SIP-HI). However, optimizing SFMs for SIP-HI has been…

Artificial Intelligence · Computer Science 2025-05-14 Haoshuai Zhou , Boxuan Cao , Changgeng Mo , Linkai Li , Shan Xiang Wang

Current multichannel speech enhancement algorithms typically assume a stationary sound source, a common mismatch with reality that limits their performance in real-world scenarios. This paper focuses on attention-driven spatial filtering…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-19 Yuzhu Wang , Archontis Politis , Tuomas Virtanen

When the parameters of Bayesian Short-time Spectral Amplitude (STSA) estimator for speech enhancement are selected based on the characteristics of the human auditory system, the gain function of the estimator becomes more flexible. Although…

Sound · Computer Science 2025-12-18 Suman Samui

This paper presents a self-supervised learning framework, named MGF, for general-purpose speech representation learning. In the design of MGF, speech hierarchy is taken into consideration. Specifically, we propose to use generative learning…

Sound · Computer Science 2021-02-04 Yucheng Zhao , Dacheng Yin , Chong Luo , Zhiyuan Zhao , Chuanxin Tang , Wenjun Zeng , Zheng-Jun Zha