中文
相关论文

相关论文: On Front-end Gain Invariant Modeling for Wake Word…

200 篇论文

Deep neural network models for speech recognition have achieved great success recently, but they can learn incorrect associations between the target and nuisance factors of speech (e.g., speaker identities, background noise, etc.), which…

计算与语言 · 计算机科学 2019-07-09 I-Hung Hsu , Ayush Jaiswal , Premkumar Natarajan

Single-channel speech enhancement approaches do not always improve automatic recognition rates in the presence of noise, because they can introduce distortions unhelpful for recognition. Following a trend towards end-to-end training of…

声音 · 计算机科学 2021-12-14 Peter Plantinga , Deblin Bagchi , Eric Fosler-Lussier

Recently, Over-the-Air (OTA) computation has emerged as a promising federated learning (FL) paradigm that leverages the waveform superposition properties of the wireless channel to realize fast model updates. Prior work focused on the OTA…

机器学习 · 计算机科学 2024-04-01 Muhammad Faraz Ul Abrar , Nicolò Michelusi

Neural latent variable models enable the discovery of interesting structure in speech audio data. This paper presents a comparison of two different approaches which are broadly based on predicting future time-steps or auto-encoding the…

音频与语音处理 · 电气工程与系统科学 2020-10-28 Henry Zhou , Alexei Baevski , Michael Auli

Reliable fault detection is an essential requirement for safe and efficient operation of complex mechanical systems in various industrial applications. Despite the abundance of existing approaches and the maturity of the fault detection…

信号处理 · 电气工程与系统科学 2024-08-19 Tianfu Li , Chuang Sun , Ruqiang Yan , Xuefeng Chen

This study tackles unsupervised subword modeling in the zero-resource scenario, learning frame-level speech representation that is phonetically discriminative and speaker-invariant, using only untranscribed speech for target languages.…

音频与语音处理 · 电气工程与系统科学 2020-10-29 Siyuan Feng , Tan Lee

We introduce EffiFusion-GAN (Efficient Fusion Generative Adversarial Network), a lightweight yet powerful model for speech enhancement. The model integrates depthwise separable convolutions within a multi-scale block to capture diverse…

声音 · 计算机科学 2025-08-21 Bin Wen , Tien-Ping Tan

With recent advances of diffusion model, generative speech enhancement (SE) has attracted a surge of research interest due to its great potential for unseen testing noises. However, existing efforts mainly focus on inherent properties of…

音频与语音处理 · 电气工程与系统科学 2024-06-05 Yuchen Hu , Chen Chen , Ruizhe Li , Qiushi Zhu , Eng Siong Chng

Precise forecasting of significant wave height (Hs) is essential for the development and utilization of wave energy. The challenges in predicting Hs arise from its non-linear and non-stationary characteristics. The combination of…

机器学习 · 计算机科学 2025-05-13 Jianxin Zhang , Lianzi Jiang , Xinyu Han , Xiangrong Wang

Integrating front-end speech enhancement (SE) models with self-supervised learning (SSL)-based speech models is effective for downstream tasks in noisy conditions. SE models are commonly fine-tuned using SSL representations with mean…

计算与语言 · 计算机科学 2026-01-30 Amit Meghanani , Thomas Hain

This paper examines the integration of real-time talking-head generation for interviewer training, focusing on overcoming challenges in Audio Feature Extraction (AFE), which often introduces latency and limits responsiveness in real-time…

In a typical sound event detection (SED) system, the existence of a sound event is detected at a frame level, and consecutive frames with the same event detected are combined as one sound event. The median filter is applied as a…

声音 · 计算机科学 2024-03-21 Tao Song

This work focuses on the validation of the dynamic wake meandering (DWM) model against large eddy simulation (LES). The wake deficit, mean deflection, and meandering under different wind turbine misalignment angles in yaw and tilt, for the…

流体动力学 · 物理学 2023-08-03 Irene Rivera-Arreba , Zhaobin Li , Xiaolei Yang , Erin E. Bachynski-Polić

Identifying user-defined keywords is crucial for personalizing interactions with smart devices. Previous approaches of user-defined keyword spotting (UDKWS) have relied on short-term spectral features such as mel frequency cepstral…

声音 · 计算机科学 2024-05-24 Kesavaraj V , Anuprabha M , Anil Kumar Vuppala

Sampling rate offsets (SROs) between devices in a heterogeneous wireless acoustic sensor network (WASN) can hinder the ability of distributed adaptive algorithms to perform as intended when they rely on coherent signal processing. In this…

音频与语音处理 · 电气工程与系统科学 2023-02-13 Paul Didier , Toon van Waterschoot , Simon Doclo , Marc Moonen

Incipient fault detection in power distribution systems is crucial to improve the reliability of the grid. However, the non-stationary nature and the inadequacy of the training dataset due to the self-recovery of the incipient fault signal,…

信号处理 · 电气工程与系统科学 2023-02-21 Qiyue Li , Huan Luo , Hong Cheng , Yuxing Deng , Wei Sun , Weitao Li , Zhi Liu

This article proposes a robust brain-inspired audio feature extractor (RBA-FE) model for depression diagnosis, using an improved hierarchical network architecture. Most deep learning models achieve state-of-the-art performance for…

声音 · 计算机科学 2025-06-10 Yu-Xuan Wu , Ziyan Huang , Bin Hu , Zhi-Hong Guan

In previous work, we proposed a variational autoencoder-based (VAE) Bayesian permutation training speech enhancement (SE) method (PVAE) which indicated that the SE performance of the traditional deep neural network-based (DNN) method could…

音频与语音处理 · 电气工程与系统科学 2022-05-12 Yang Xiang , Jesper Lisby Højvang , Morten Højfeldt Rasmussen , Mads Græsbøll Christensen

This paper presents a generative approach to speech enhancement based on a recurrent variational autoencoder (RVAE). The deep generative speech model is trained using clean speech signals only, and it is combined with a nonnegative matrix…

机器学习 · 计算机科学 2020-02-11 Simon Leglaive , Xavier Alameda-Pineda , Laurent Girin , Radu Horaud

This work focuses on online dereverberation for hearing devices using the weighted prediction error (WPE) algorithm. WPE filtering requires an estimate of the target speech power spectral density (PSD). Recently deep neural networks (DNNs)…

音频与语音处理 · 电气工程与系统科学 2022-05-06 Jean-Marie Lemercier , Joachim Thiemann , Raphael Koning , Timo Gerkmann