English
Related papers

Related papers: Deep Learning-Based Joint Control of Acoustic Echo…

200 papers

Data-driven models achieve successful results in Speech Emotion Recognition (SER). However, these models, which are often based on general acoustic features or end-to-end approaches, show poor performance when the testing set has a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-15 Duowei Tang , Peter Kuppens , Lucca Geurts , Toon van Waterschoot

This paper presents a real-time Acoustic Echo Cancellation (AEC) algorithm submitted to the AEC-Challenge. The algorithm consists of three modules: Generalized Cross-Correlation with PHAse Transform (GCC-PHAT) based time delay compensation,…

Sound · Computer Science 2021-02-19 Ziteng Wang , Yueyue Na , Zhang Liu , Biao Tian , Qiang Fu

In reverberant conditions with a single speaker, each far-field microphone records a reverberant version of the same speaker signal at a different location. In over-determined conditions, where there are multiple microphones but only one…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-14 Zhong-Qiu Wang

Recently, variational autoencoder (VAE), a deep representation learning (DRL) model, has been used to perform speech enhancement (SE). However, to the best of our knowledge, current VAE-based SE methods only apply VAE to the model speech…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-25 Yang Xiang , Jesper Lisby Højvang , Morten Højfeldt Rasmussen , Mads Græsbøll Christensen

Two algorithms for combined acoustic echo cancellation (AEC) and noise reduction (NR) are analysed, namely the generalised echo and interference canceller (GEIC) and the extended multichannel Wiener filter (MWFext). Previously, these…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-15 Arnout Roebben , Toon van Waterschoot , Marc Moonen

The performance of speech enhancement algorithms in a multi-speaker scenario depends on correctly identifying the target speaker to be enhanced. Auditory attention decoding (AAD) methods allow to identify the target speaker which the…

Sound · Computer Science 2020-05-12 Ali Aroudi , Marc Delcroix , Tomohiro Nakatani , Keisuke Kinoshita , Shoko Araki , Simon Doclo

Deep neural networks (DNNs) used for brain-computer-interface (BCI) classification are commonly expected to learn general features when trained across a variety of contexts, such that these features could be fine-tuned to specific contexts.…

Machine Learning · Computer Science 2021-01-29 Demetres Kostas , Stephane Aroca-Ouellette , Frank Rudzicz

Speech enhancement and source localization has been active research for several decades with a wide range of real-world applications. Recently, the Deep Complex Convolution Recurrent network (DCCRN) has yielded impressive enhancement…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-22 Yuan Chen , Yicheng Hsu , Mingsian R. Bai

The state-of-art models for speech synthesis and voice conversion are capable of generating synthetic speech that is perceptually indistinguishable from bonafide human speech. These methods represent a threat to the automatic speaker…

Machine Learning · Computer Science 2019-07-11 Moustafa Alzantot , Ziqi Wang , Mani B. Srivastava

This paper introduces an explainable DNN-based beamformer with a postfilter (ExNet-BF+PF) for multichannel signal processing. Our approach combines the U-Net network with a beamformer structure to address this problem. The method involves a…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-19 Adi Cohen , Daniel Wong , Jung-Suk Lee , Sharon Gannot

With the increasing demand for audio communication and online conference, ensuring the robustness of Acoustic Echo Cancellation (AEC) under the complicated acoustic scenario including noise, reverberation and nonlinear distortion has become…

Sound · Computer Science 2022-02-16 Shimin Zhang , Yuxiang Kong , Shubo Lv , Yanxin Hu , Lei Xie

Monaural speech dereverberation is a very challenging task because no spatial cues can be used. When the additive noises exist, this task becomes more challenging. In this paper, we propose a joint training method for simultaneous speech…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-07 Cunhang Fan , Jianhua Tao , Bin Liu , Jiangyan Yi , Zhengqi Wen

Recently, our proposed recurrent neural network (RNN) based all deep learning minimum variance distortionless response (ADL-MVDR) beamformer method yielded superior performance over the conventional MVDR by replacing the matrix inversion…

Sound · Computer Science 2021-04-27 Xiyun Li , Yong Xu , Meng Yu , Shi-Xiong Zhang , Jiaming Xu , Bo Xu , Dong Yu

Children speech recognition is indispensable but challenging due to the diversity of children's speech. In this paper, we propose a filter-based discriminative autoencoder for acoustic modeling. To filter out the influence of various…

Computation and Language · Computer Science 2022-05-24 Chiang-Lin Tai , Hung-Shin Lee , Yu Tsao , Hsin-Min Wang

To achieve robust far-field automatic speech recognition (ASR), existing techniques typically employ an acoustic front end (AFE) cascaded with a neural transducer (NT) ASR model. The AFE output, however, could be unreliable, as the…

In many speech recording applications, the recorded desired speech is corrupted by both noise and acoustic echo, such that combined noise reduction (NR) and acoustic echo cancellation (AEC) is called for. A common cascaded design…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-06 Arnout Roebben , Toon van Waterschoot , Marc Moonen

Many purely neural network based speech separation approaches have been proposed to improve objective assessment scores, but they often introduce nonlinear distortions that are harmful to modern automatic speech recognition (ASR) systems.…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-17 Zhuohuang Zhang , Yong Xu , Meng Yu , Shi-Xiong Zhang , Lianwu Chen , Donald S. Williamson , Dong Yu

This paper presents an acoustic echo canceler based on a U-Net convolutional neural network for single-talk and double-talk scenarios. U-Net networks have previously been used in the audio processing area for source separation problems…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-21 J. Silva-Rodríguez , M. F. Dolz , M. Ferrer , A. Castelló , V. Naranjo , G. Piñero

We propose a deep beamforming framework for enhancing target speaker(s) in multi-speaker environments. A deep neural network (DNN) is trained to estimate beamforming weights directly from noisy multichannel inputs while satisfying linear…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-21 Ilai Zaidel , Ori Engel , Bar Engel , Sharon Gannot

Although the conventional mask-based minimum variance distortionless response (MVDR) could reduce the non-linear distortion, the residual noise level of the MVDR separated speech is still high. In this paper, we propose a spatio-temporal…

Sound · Computer Science 2021-04-06 Yong Xu , Zhuohuang Zhang , Meng Yu , Shi-Xiong Zhang , Dong Yu