中文
相关论文

相关论文: Differentiable Time-Varying IIR Filtering for Real…

200 篇论文

In this paper, we explore a continuous modeling approach for deep-learning-based speech enhancement, focusing on the denoising process. We use a state variable to indicate the denoising process. The starting state is noisy speech and the…

音频与语音处理 · 电气工程与系统科学 2024-01-09 Zilu Guo , Jun Du , CHin-Hui Lee

Speech Emotion Recognition (SER) systems often degrade in performance when exposed to the unpredictable acoustic interference found in real-world environments. Additionally, the opacity of deep learning models hinders their adoption in…

声音 · 计算机科学 2025-12-23 Sudip Chakrabarty , Pappu Bishwas , Rajdeep Chatterjee

Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hindered their applications to speech synthesis. This paper…

音频与语音处理 · 电气工程与系统科学 2022-04-22 Rongjie Huang , Max W. Y. Lam , Jun Wang , Dan Su , Dong Yu , Yi Ren , Zhou Zhao

Most speech enhancement algorithms make use of the short-time Fourier transform (STFT), which is a simple and flexible time-frequency decomposition that estimates the short-time spectrum of a signal. However, the duration of short STFT…

声音 · 计算机科学 2015-09-03 Scott Wisdom , Thomas Powers , Les Atlas , James Pitton

While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding, they frequently falter in fine-grained perception tasks that require identifying tiny objects or discerning subtle…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Jilong Zhu , Yang Feng

An integrate-and-fire time-encoding machine (IF-TEM) is an effective asynchronous sampler that translates amplitude information into non-uniform time sequences. In this work, we propose a novel Adaptive IF-TEM (AIF-TEM) approach. This…

信号处理 · 电气工程与系统科学 2026-03-18 Aseel Omar , Alejandro Cohen

Transient loud intrusions, often occurring in noisy environments, can completely overpower speech signal and lead to an inevitable loss of information. While existing algorithms for noise suppression can yield impressive results, their…

声音 · 计算机科学 2020-11-12 Mikolaj Kegler , Pierre Beckmann , Milos Cernak

Mel-scale spectrum features are used in various recognition and classification tasks on speech signals. There is no reason to expect that these features are optimal for all different tasks, including speaker verification (SV). This paper…

音频与语音处理 · 电气工程与系统科学 2022-06-16 Jingyu Li , Yusheng Tian , Tan Lee

Diffusion and flow-based generative models have shown strong potential for image restoration. However, image denoising under unknown and varying noise conditions remains challenging, because the learned vector fields may become inconsistent…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Jigang Duan , Genwei Ma , Xu Jiang , Wenfeng Xu , Ping Yang , Xing Zhao

A flexible discriminative image denoiser is introduced in which multi-task learning methods are applied to a densoising FCN based on U-Net. The activations of the U-Net model are modified by affine transforms that are a learned function of…

图像与视频处理 · 电气工程与系统科学 2020-11-26 Anthony Kelly

Dynamic graph signal processing provides a principled framework for analyzing time-varying data defined on irregular graph domains. However, existing joint time-vertex transforms such as the joint time-vertex fractional Fourier transform…

信号处理 · 电气工程与系统科学 2025-11-21 Manjun Cui , Ziqi Yan , Yangfan He , Zhichao Zhang

Finance is a particularly challenging application area for deep learning models due to low noise-to-signal ratio, non-stationarity, and partial observability. Non-deliverable-forwards (NDF), a derivatives contract used in foreign exchange…

机器学习 · 计算机科学 2019-09-25 Michael Poli , Jinkyoo Park , Ilija Ilievski

Federated learning (FL) scenarios inherently generate a large communication overhead by frequently transmitting neural network updates between clients and server. To minimize the communication cost, introducing sparsity in conjunction with…

This work proposes a neural network to extensively exploit spatial information for multichannel joint speech separation, denoising and dereverberation, named SpatialNet. In the short-time Fourier transform (STFT) domain, the proposed…

声音 · 计算机科学 2023-12-25 Changsheng Quan , Xiaofei Li

Diffusion models are the go-to method for Text-to-Image generation, but their iterative denoising processes has high inference latency. Quantization reduces compute time by using lower bitwidths, but applies a fixed precision across all…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Basile Lewandowski , Simon Kurz , Aditya Shankar , Robert Birke , Jian-Jia Chen , Lydia Y. Chen

Leveraging the fact that speaker identity and content vary on different time scales, \acrlong{fhvae} (\acrshort{fhvae}) uses different latent variables to symbolize these two attributes. Disentanglement of these attributes is carried out by…

音频与语音处理 · 电气工程与系统科学 2023-06-16 Yuying Xie , Thomas Arildsen , Zheng-Hua Tan

With recent deep learning based approaches showing promising results in removing noise from images, the best denoising performance has been reported in a supervised learning setup that requires a large set of paired noisy images and ground…

图像与视频处理 · 电气工程与系统科学 2022-09-20 Rihuan Ke

Time-frequency analysis for non-linear and non-stationary signals is extraordinarily challenging. To capture features in these signals, it is necessary for the analysis methods to be local, adaptive and stable. In recent years,…

数值分析 · 数学 2015-10-26 Antonio Cicone , Jingfang Liu , Haomin Zhou

Noise-robust automatic speech recognition (ASR) has been commonly addressed by applying speech enhancement (SE) at the waveform level before recognition. However, speech-level enhancement does not always translate into consistent…

音频与语音处理 · 电气工程与系统科学 2026-01-09 Da-Hee Yang , Joon-Hyuk Chang

Recent advances in self-supervised learning (SSL) on Transformers have significantly improved speaker verification (SV) by providing domain-general speech representations. However, existing approaches have underutilized the multi-layered…

音频与语音处理 · 电气工程与系统科学 2025-12-16 Jin Sob Kim , Hyun Joon Park , Wooseok Shin , Juan Yun , Sung Won Han