English
Related papers

Related papers: Impulse Response Data Augmentation and Deep Neural…

200 papers

In recent years, dynamic parameterization of acoustic environments has garnered attention in audio processing. This focus includes room volume and reverberation time (RT60), which define local acoustics independent of sound source and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-10 Chunxi Wang , Maoshen Jia , Meiran Li , Changchun Bao , Wenyu Jin

The reverberation time is one of the most important parameters used to characterize the acoustic property of an enclosure. In real-world scenarios, it is much more convenient to estimate the reverberation time blindly from recorded speech…

Sound · Computer Science 2021-12-10 Kaitong Zheng , Chengshi Zheng , Jinqiu Sang , Yulong Zhang , Xiaodong Li

In IEEE 802.11bc, the broadcast mode on wireless local area networks (WLANs), data rate control that is based on acknowledgement (ACK) mechanism similar to the one in the current IEEE 802.11 WLANs is not applicable because ACK mechanism is…

Networking and Internet Architecture · Computer Science 2022-03-04 T. Kanda , Y. Koda , K. Yamamoto , T. Nishio

This paper presents a speech intelligibility model based on automatic speech recognition (ASR), combining phoneme probabilities from deep neural networks (DNN) and a performance measure that estimates the word error rate from these…

We propose a neural network model that can separate target speech sources from interfering sources at different angular regions using two microphones. The model is trained with simulated room impulse responses (RIRs) using omni-directional…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-18 Yang Yang , George Sung , Shao-Fu Shih , Hakan Erdogan , Chehung Lee , Matthias Grundmann

In recent years, machine learning approaches to modelling guitar amplifiers and effects pedals have been widely investigated and have become standard practice in some consumer products. In particular, recurrent neural networks (RNNs) are a…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-11 Alistair Carson , Alec Wright , Jatin Chowdhury , Vesa Välimäki , Stefan Bilbao

While machine learning techniques are traditionally resource intensive, we are currently witnessing an increased interest in hardware and energy efficient approaches. This need for resource-efficient machine learning is primarily driven by…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-23 Lukas Pfeifenberger , Matthias Zöhrer , Günther Schindler , Wolfgang Roth , Holger Fröning , Franz Pernkopf

In this paper, we consider the importance of channel measurement data from specific sites and its impact on air interface optimization and test. Currently, a range of statistical channel models including 3GPP 38.901 tapped delay line (TDL),…

Signal Processing · Electrical Eng. & Systems 2024-10-28 Johnathan Corgan , Nitin Nair , Rajib Bhattacharjea , Wan Liu , Serhat Tadik , Tom Tsou , Timothy J. O'Shea

Early detection and treatment of depression is essential in promoting remission, preventing relapse, and reducing the emotional burden of the disease. Current diagnoses are primarily subjective, inconsistent across professionals, and…

Machine Learning · Computer Science 2020-02-03 Karol Chlasta , Krzysztof Wołk , Izabela Krejtz

Room acoustical parameters (RAPs), room geometrical parameters (RGPs) and instantaneous occupancy level are essential metrics for parameterizing the room acoustical characteristics (RACs) of a sound field around a listener's local…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-02 Lijun Wang , Yixian Lu , Ziyan Gao , Kai Li , Jianqiang Huang , Yuntao Kong , Shogo Okada

Non-data-aided (NDA) parameter estimation is considered for binary-phase-shift-keying transmission in an additive white Gaussian noise channel. Cramer-Rao lower bounds (CRLBs) for signal amplitude, noise variance, channel reliability…

Information Theory · Computer Science 2007-07-13 Fredrik Brannstrom , Lars K. Rasmussen

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of…

Sound · Computer Science 2019-04-03 Hyeong-Seok Choi , Jang-Hyun Kim , Jaesung Huh , Adrian Kim , Jung-Woo Ha , Kyogu Lee

Indoor localization is a challenging problem that - unlike outdoor localization - lacks a universal and robust solution. Machine Learning (ML), particularly Deep Learning (DL), methods have been investigated as a promising approach.…

Systems and Control · Electrical Eng. & Systems 2024-08-29 Omer Gokalp Serbetci , Daoud Burghal , Andreas F. Molisch

Today, the optimal performance of existing noise-suppression algorithms, both data-driven and those based on classic statistical methods, is range bound to specific levels of instantaneous input signal-to-noise ratios. In this paper, we…

Machine Learning · Computer Science 2018-07-30 Rasool Fakoor , Xiaodong He , Ivan Tashev , Shuayb Zarar

We introduce Whisper-RIR-Mega, a benchmark dataset of paired clean and reverberant speech for evaluating automatic speech recognition (ASR) robustness to room acoustics. Each sample pairs a clean LibriSpeech utterance with the same…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-17 Mandip Goswami

We propose ARiSE, an auto-regressive algorithm for multi-channel speech enhancement. ARiSE improves existing deep neural network (DNN) based frame-online multi-channel speech enhancement models by introducing auto-regressive connections,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-09 Pengjie Shen , Xueliang Zhang , Zhong-Qiu Wang

Whisper's robust performance in automatic speech recognition (ASR) is often attributed to its massive 680k-hour training set, an impractical scale for most researchers. In this work, we examine how linguistic and acoustic diversity in…

Computation and Language · Computer Science 2025-05-28 Dancheng Liu , Amir Nassereldine , Chenhui Xu , Jinjun Xiong

Deep learning technology has been widely applied to speech enhancement. While testing the effectiveness of various network structures, researchers are also exploring the improvement of the loss function used in network training. Although…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-25 Tianrui Wang , Weibin Zhu

This work introduces a robust single-channel inverse filter for dereverberation of non-ideal recordings, validated on real audio. The developed method focuses on the calculation and modification of a discrete impulse response in order to…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-13 Stefan Ciba

This paper proposes a practical approach to estimate the direct-to-reverberant energy ratio (DRR) using a spherical microphone array without having knowledge of the source signal. We base our estimation on a theoretical relationship between…

Sound · Computer Science 2015-11-02 Hanchi Chen , Prasanga N. Samarasinghe , Thushara D. Abhayapala , Wen Zhang