English
Related papers

Related papers: On Front-end Gain Invariant Modeling for Wake Word…

200 papers

Deep neural network models for speech recognition have achieved great success recently, but they can learn incorrect associations between the target and nuisance factors of speech (e.g., speaker identities, background noise, etc.), which…

Computation and Language · Computer Science 2019-07-09 I-Hung Hsu , Ayush Jaiswal , Premkumar Natarajan

Single-channel speech enhancement approaches do not always improve automatic recognition rates in the presence of noise, because they can introduce distortions unhelpful for recognition. Following a trend towards end-to-end training of…

Sound · Computer Science 2021-12-14 Peter Plantinga , Deblin Bagchi , Eric Fosler-Lussier

Recently, Over-the-Air (OTA) computation has emerged as a promising federated learning (FL) paradigm that leverages the waveform superposition properties of the wireless channel to realize fast model updates. Prior work focused on the OTA…

Machine Learning · Computer Science 2024-04-01 Muhammad Faraz Ul Abrar , Nicolò Michelusi

Neural latent variable models enable the discovery of interesting structure in speech audio data. This paper presents a comparison of two different approaches which are broadly based on predicting future time-steps or auto-encoding the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-28 Henry Zhou , Alexei Baevski , Michael Auli

Reliable fault detection is an essential requirement for safe and efficient operation of complex mechanical systems in various industrial applications. Despite the abundance of existing approaches and the maturity of the fault detection…

Signal Processing · Electrical Eng. & Systems 2024-08-19 Tianfu Li , Chuang Sun , Ruqiang Yan , Xuefeng Chen

This study tackles unsupervised subword modeling in the zero-resource scenario, learning frame-level speech representation that is phonetically discriminative and speaker-invariant, using only untranscribed speech for target languages.…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-29 Siyuan Feng , Tan Lee

We introduce EffiFusion-GAN (Efficient Fusion Generative Adversarial Network), a lightweight yet powerful model for speech enhancement. The model integrates depthwise separable convolutions within a multi-scale block to capture diverse…

Sound · Computer Science 2025-08-21 Bin Wen , Tien-Ping Tan

With recent advances of diffusion model, generative speech enhancement (SE) has attracted a surge of research interest due to its great potential for unseen testing noises. However, existing efforts mainly focus on inherent properties of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-05 Yuchen Hu , Chen Chen , Ruizhe Li , Qiushi Zhu , Eng Siong Chng

Precise forecasting of significant wave height (Hs) is essential for the development and utilization of wave energy. The challenges in predicting Hs arise from its non-linear and non-stationary characteristics. The combination of…

Machine Learning · Computer Science 2025-05-13 Jianxin Zhang , Lianzi Jiang , Xinyu Han , Xiangrong Wang

Integrating front-end speech enhancement (SE) models with self-supervised learning (SSL)-based speech models is effective for downstream tasks in noisy conditions. SE models are commonly fine-tuned using SSL representations with mean…

Computation and Language · Computer Science 2026-01-30 Amit Meghanani , Thomas Hain

This paper examines the integration of real-time talking-head generation for interviewer training, focusing on overcoming challenges in Audio Feature Extraction (AFE), which often introduces latency and limits responsiveness in real-time…

In a typical sound event detection (SED) system, the existence of a sound event is detected at a frame level, and consecutive frames with the same event detected are combined as one sound event. The median filter is applied as a…

Sound · Computer Science 2024-03-21 Tao Song

This work focuses on the validation of the dynamic wake meandering (DWM) model against large eddy simulation (LES). The wake deficit, mean deflection, and meandering under different wind turbine misalignment angles in yaw and tilt, for the…

Fluid Dynamics · Physics 2023-08-03 Irene Rivera-Arreba , Zhaobin Li , Xiaolei Yang , Erin E. Bachynski-Polić

Identifying user-defined keywords is crucial for personalizing interactions with smart devices. Previous approaches of user-defined keyword spotting (UDKWS) have relied on short-term spectral features such as mel frequency cepstral…

Sound · Computer Science 2024-05-24 Kesavaraj V , Anuprabha M , Anil Kumar Vuppala

Sampling rate offsets (SROs) between devices in a heterogeneous wireless acoustic sensor network (WASN) can hinder the ability of distributed adaptive algorithms to perform as intended when they rely on coherent signal processing. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-13 Paul Didier , Toon van Waterschoot , Simon Doclo , Marc Moonen

Incipient fault detection in power distribution systems is crucial to improve the reliability of the grid. However, the non-stationary nature and the inadequacy of the training dataset due to the self-recovery of the incipient fault signal,…

Signal Processing · Electrical Eng. & Systems 2023-02-21 Qiyue Li , Huan Luo , Hong Cheng , Yuxing Deng , Wei Sun , Weitao Li , Zhi Liu

This article proposes a robust brain-inspired audio feature extractor (RBA-FE) model for depression diagnosis, using an improved hierarchical network architecture. Most deep learning models achieve state-of-the-art performance for…

Sound · Computer Science 2025-06-10 Yu-Xuan Wu , Ziyan Huang , Bin Hu , Zhi-Hong Guan

In previous work, we proposed a variational autoencoder-based (VAE) Bayesian permutation training speech enhancement (SE) method (PVAE) which indicated that the SE performance of the traditional deep neural network-based (DNN) method could…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-12 Yang Xiang , Jesper Lisby Højvang , Morten Højfeldt Rasmussen , Mads Græsbøll Christensen

This paper presents a generative approach to speech enhancement based on a recurrent variational autoencoder (RVAE). The deep generative speech model is trained using clean speech signals only, and it is combined with a nonnegative matrix…

Machine Learning · Computer Science 2020-02-11 Simon Leglaive , Xavier Alameda-Pineda , Laurent Girin , Radu Horaud

This work focuses on online dereverberation for hearing devices using the weighted prediction error (WPE) algorithm. WPE filtering requires an estimate of the target speech power spectral density (PSD). Recently deep neural networks (DNNs)…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-06 Jean-Marie Lemercier , Joachim Thiemann , Raphael Koning , Timo Gerkmann