English
Related papers

Related papers: Single-step Controllable Music Bandwidth Extension…

200 papers

The direct expansion of deep neural network (DNN) based wide-band speech enhancement (SE) to full-band processing faces the challenge of low frequency resolution in low frequency range, which would highly likely lead to deteriorated…

Sound · Computer Science 2022-06-28 Zhongshu Hou , Qinwen Hu , Kai Chen , Jing Lu

Diffusion models have emerged as powerful tools for generative tasks, producing high-quality outputs across diverse domains. However, how the generated data responds to the initial noise perturbation in diffusion models remains…

Machine Learning · Computer Science 2025-02-10 Bowen Song , Zecheng Zhang , Zhaoxu Luo , Jason Hu , Wei Yuan , Jing Jia , Zhengxu Tang , Guanyang Wang , Liyue Shen

Diffusion models have shown a great ability at bridging the performance gap between predictive and generative approaches for speech enhancement. We have shown that they may even outperform their predictive counterparts for non-additive…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-13 Jean-Marie Lemercier , Julius Richter , Simon Welker , Timo Gerkmann

The two-dimensional backward-facing step flow is a canonical example of noise amplifier flow: global linear stability analysis predicts that it is stable, but perturbations can undergo large amplification in space and time as a result of…

Fluid Dynamics · Physics 2014-12-05 Edouard Boujo , François Gallaire

Diffusion models trained on noisy datasets often reproduce high-frequency training artifacts, significantly degrading generation quality. To address this, we propose SCoRe (Spectral Cutoff Regeneration), a training-free, generation-time…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Yuta Matsuzaki , Seiichi Uchida , Shumpei Takezaki

Convolution and cross-correlation are the basis of filtering and pattern or template matching in multimedia signal processing. We propose two throughput scaling options for any one-dimensional convolution kernel in programmable processors…

Multimedia · Computer Science 2012-01-17 Mohammad Ashraful Anam , Yiannis Andreopoulos

To improve the performance in identifying the faults under strong noise for rotating machinery, this paper presents a dynamic feature reconstruction signal graph method, which plays the key role of the proposed end-to-end fault diagnosis…

Signal Processing · Electrical Eng. & Systems 2023-10-02 Wenbin He , Jianxu Mao , Yaonan Wang , Zhe Li , Qiu Fang , Haotian Wu

The restoration of nonlinearly distorted audio signals, alongside the identification of the applied memoryless nonlinear operation, is studied. The paper focuses on the difficult but practically important case in which both the nonlinearity…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-13 Michal Švento , Eloi Moliner , Lauri Juvela , Alec Wright , Vesa Välimäki

Audio bandwidth extension involves the realistic reconstruction of high-frequency spectra from bandlimited observations. In cases where the lowpass degradation is unknown, such as in restoring historical audio recordings, this becomes a…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-31 Eloi Moliner , Filip Elvander , Vesa Välimäki

Diffusion models have been shown to implicitly generate visual content autoregressively in the frequency domain, where low-frequency components are generated earlier in the denoising process while high-frequency details emerge only in later…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Howard Xiao , Brian Chao , Lior Yariv , Gordon Wetzstein

Noise reduction techniques based on deep learning have demonstrated impressive performance in enhancing the overall quality of recorded speech. While these approaches are highly performant, their application in audio engineering can be…

Sound · Computer Science 2023-10-18 Christian J. Steinmetz , Thomas Walther , Joshua D. Reiss

Versatile audio super-resolution (SR) is the challenging task of restoring high-frequency components from low-resolution audio with sampling rates between 4kHz and 32kHz in various domains such as music, speech, and sound effects. Previous…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-22 Jaekwon Im , Juhan Nam

For music indexing robust to sound degradations and scalable for big music catalogs, this scientific report presents an approach based on audio descriptors relevant to the music content and invariant to sound transformations (noise…

Signal Processing · Electrical Eng. & Systems 2024-03-04 Rémi Mignot , Geoffroy Peeters

Diffusion posterior sampling solves inverse problems by combining a pretrained diffusion prior with measurement-consistency guidance, but it often fails to recover fine details because measurement terms are applied in a manner that is…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Feng Tian , Yixuan Li , Weili Zeng , Weitian Zhang , Yichao Yan , Xiaokang Yang

Multimodal contrastive models have achieved strong performance in text-audio retrieval and zero-shot settings, but improving joint embedding spaces remains an active research area. Less attention has been given to making these systems…

Sound · Computer Science 2025-06-25 Julien Guinot , Elio Quinton , György Fazekas

Image restoration aims to recover high-quality images from degraded observations. When the degradation process is known, the recovery problem can be formulated as an inverse problem, and in a Bayesian context, the goal is to sample a clean…

Image and Video Processing · Electrical Eng. & Systems 2025-10-13 Darshan Thaker , Abhishek Goyal , René Vidal

Deepfake (DF) audio detectors still struggle to generalize to out of distribution inputs. A central reason is spectral bias, the tendency of neural networks to learn low-frequency structure before high-frequency (HF) details, which both…

Sound · Computer Science 2025-11-27 Ido Nitzan HIdekel , Gal lifshitz , Khen Cohen , Dan Raviv

Semantic Communication (SC) is an emerging technology that has attracted much attention in the sixth-generation (6G) mobile communication systems. However, few literature has fully considered the perceptual quality of the reconstructed…

Image and Video Processing · Electrical Eng. & Systems 2024-10-04 Kexin Zhang , Lixin Li , Wensheng Lin , Yuna Yan , Wenchi Cheng , Zhu Han

Speech restoration in real-world conditions is challenging due to compounded distortions and mismatches between input and desired output rates. Most existing systems assume a fixed and shared input-output rate, relying on external…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-29 Ui-Hyeop Shin , Jaehyun Ko , Woocheol Jeong , Hyung-Min Park

Denoising diffusion probabilistic models (DDPMs) can be utilized to recover a clean signal from its degraded observation(s) by conditioning the model on the degraded signal. The degraded signals are themselves contaminated versions of the…

Image and Video Processing · Electrical Eng. & Systems 2025-06-10 Ching-Hua Lee , Chouchang Yang , Jaejin Cho , Yashas Malur Saidutta , Rakshith Sharma Srinivasa , Yilin Shen , Hongxia Jin