中文
相关论文

相关论文: Direction-Aware Adaptive Online Neural Speech Enha…

200 篇论文

We propose ARiSE, an auto-regressive algorithm for multi-channel speech enhancement. ARiSE improves existing deep neural network (DNN) based frame-online multi-channel speech enhancement models by introducing auto-regressive connections,…

音频与语音处理 · 电气工程与系统科学 2025-06-09 Pengjie Shen , Xueliang Zhang , Zhong-Qiu Wang

Multi-channel speech separation using speaker's directional information has demonstrated significant gains over blind speech separation. However, it has two limitations. First, substantial performance degradation is observed when the coming…

声音 · 计算机科学 2023-02-28 Rongzhi Gu , Shi-Xiong Zhang , Dong Yu

Acoustic beamformers have been widely used to enhance audio signals. Currently, the best methods are the deep neural network (DNN)-powered variants of the generalized eigenvalue and minimum-variance distortionless response beamformers and…

音频与语音处理 · 电气工程与系统科学 2020-08-12 Yuichiro Koyama , Bhiksha Raj

Speech separation algorithms are often used to separate the target speech from other interfering sources. However, purely neural network based speech separation systems often cause nonlinear distortion that is harmful for automatic speech…

音频与语音处理 · 电气工程与系统科学 2021-02-10 Zhuohuang Zhang , Yong Xu , Meng Yu , Shi-Xiong Zhang , Lianwu Chen , Dong Yu

In reverberant conditions with a single speaker, each far-field microphone records a reverberant version of the same speaker signal at a different location. In over-determined conditions, where there are multiple microphones but only one…

音频与语音处理 · 电气工程与系统科学 2024-08-14 Zhong-Qiu Wang

Multi-channel speech enhancement with ad-hoc sensors has been a challenging task. Speech model guided beamforming algorithms are able to recover natural sounding speech, but the speech models tend to be oversimplified or the inference would…

计算与语言 · 计算机科学 2018-02-16 Kaizhi Qian , Yang Zhang , Shiyu Chang , Xuesong Yang , Dinei Florencio , Mark Hasegawa-Johnson

Speech enhancement in ad-hoc microphone arrays is often hindered by the asynchronization of the devices composing the microphone array. Asynchronization comes from sampling time offset and sampling rate offset which inevitably occur when…

音频与语音处理 · 电气工程与系统科学 2023-08-01 Nicolas Furnon , Romain Serizel , Slim Essid , Irina Illina

Modern smart glasses leverage advanced audio sensing and machine learning technologies to offer real-time transcribing and captioning services, considerably enriching human experiences in daily communications. However, such systems…

This study proposes a trainable adaptive window switching (AWS) method and apply it to a deep-neural-network (DNN) for speech enhancement in the modified discrete cosine transform domain. Time-frequency (T-F) mask processing in the…

音频与语音处理 · 电气工程与系统科学 2019-02-21 Yuma Koizumi , Noboru Harada , Yoichi Haneda

This paper addresses the problem of automatic speech recognition (ASR) of a target speaker in background speech. The novelty of our approach is that we focus on a wakeup keyword, which is usually used for activating ASR systems like smart…

音频与语音处理 · 电气工程与系统科学 2018-11-08 Yusuke Kida , Dung Tran , Motoi Omachi , Toru Taniguchi , Yuya Fujita

Most neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a neural source filter network for speech enhancement.…

声音 · 计算机科学 2022-10-31 Shulin He , Wei Rao , Jinjiang Liu , Jun Chen , Yukai Ju , Xueliang Zhang , Yannan Wang , Shidong Shang

Enhancing noisy speech is an important task to restore its quality and to improve its intelligibility. In traditional non-machine-learning (ML) based approaches the parameters required for noise reduction are estimated blindly from the…

声音 · 计算机科学 2018-01-16 Robert Rehr , Timo Gerkmann

In the field of speech enhancement, time domain methods have difficulties in achieving both high performance and efficiency. Recently, dual-path models have been adopted to represent long sequential features, but they still have limited…

音频与语音处理 · 电气工程与系统科学 2022-03-07 Hyun Joon Park , Byung Ha Kang , Wooseok Shin , Jin Sob Kim , Sung Won Han

This paper proposes a delayed subband LSTM network for online monaural (single-channel) speech enhancement. The proposed method is developed in the short time Fourier transform (STFT) domain. Online processing requires frame-by-frame signal…

声音 · 计算机科学 2023-12-13 Xiaofei Li , Radu Horaud

The human auditory system has the ability to selectively focus on key speech elements in an audio stream while giving secondary attention to less relevant areas such as noise or distortion within the background, dynamically adjusting its…

音频与语音处理 · 电气工程与系统科学 2026-04-09 Nursadul Mamun , John H. L. Hansen

Fixed beamforming is widely used in practice since it does not depend on the estimation of noise statistics and provides relatively stable performance. However, a single beamformer cannot adapt to varying acoustic conditions, which limits…

声音 · 计算机科学 2025-10-29 Yuanhang Qian , Kunlong Zhao , Jilu Jin , Xueqin Luo , Gongping Huang , Jingdong Chen , Jacob Benesty

Multichannel speech enhancement is widely used as a front-end in microphone array processing systems. While most existing approaches produce a single enhanced signal, direction-preserving multiple-input multiple-output (MIMO) methods…

音频与语音处理 · 电气工程与系统科学 2026-04-14 Thomas Deppisch

The joint training of speech enhancement and speaker embedding networks for speaker recognition is widely adopted under noisy acoustic environments. While effective, this paradigm often fails to leverage the generalization and robustness…

音频与语音处理 · 电气工程与系统科学 2026-04-29 Chong-Xin Gan , Peter Bell , Man-Wai Mak , Zhe Li , Zezhong Jin , Zilong Huang , Kong Aik Lee

Acoustic beamformers have been widely used to enhance audio signals. The best current methods are DNN-powered variants of the generalized eigenvalue beamformer, and DNN-based filterestimation methods that directly compute beamforming…

声音 · 计算机科学 2020-03-03 Yuichiro Koyama , Bhiksha Raj

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of…

声音 · 计算机科学 2019-04-03 Hyeong-Seok Choi , Jang-Hyun Kim , Jaesung Huh , Adrian Kim , Jung-Woo Ha , Kyogu Lee