中文
相关论文

相关论文: VRDMG: Vocal Restoration via Diffusion Posterior S…

200 篇论文

Recent advancements in solving Bayesian inverse problems have spotlighted denoising diffusion models (DDMs) as effective priors. Although these have great potential, DDM priors yield complex posterior distributions that are challenging to…

机器学习 · 统计学 2024-11-14 Yazid Janati , Badr Moufad , Alain Durmus , Eric Moulines , Jimmy Olsson

Diffusion bridge models establish probabilistic paths between arbitrary paired distributions and exhibit great potential for universal image restoration. Most existing methods merely treat them as simple variants of stochastic interpolants,…

计算机视觉与模式识别 · 计算机科学 2026-05-04 Hebaixu Wang , Jing Zhang , Haoyang Chen , Haonan Guo , Di Wang , Jiayi Ma , Bo Du

We propose DAVIS, a Diffusion-based Audio-VIsual Separation framework that solves the audio-visual sound source separation task through generative learning. Existing methods typically frame sound separation as a mask-based regression…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Chao Huang , Susan Liang , Yapeng Tian , Anurag Kumar , Chenliang Xu

In this work, we propose an approach to music source separation that uses a generative diffusion model as a last-stage refinement on top of a deterministic separator, progressively enhancing the separated sources through iterative…

声音 · 计算机科学 2026-04-28 Tornike Karchkhadze , Mohammad Rasool Izadi , Shuo Zhang , Shlomo Dubnov

Directly sending audio signals from a transmitter to a receiver across a noisy channel may absorb consistent bandwidth and be prone to errors when trying to recover the transmitted bits. On the contrary, the recent semantic communication…

声音 · 计算机科学 2023-09-15 Eleonora Grassucci , Christian Marinoni , Andrea Rodriguez , Danilo Comminiello

Despite today's prevalence of ultrasound imaging in medicine, ultrasound signal-to-noise ratio is still affected by several sources of noise and artefacts. Moreover, enhancing ultrasound image quality involves balancing concurrent factors…

图像与视频处理 · 电气工程与系统科学 2024-06-18 Yuxin Zhang , Clément Huneau , Jérôme Idier , Diana Mateus

Deep learning has shown the capability to substantially accelerate MRI reconstruction while acquiring fewer measurements. Recently, diffusion models have gained burgeoning interests as a novel group of deep learning-based generative…

图像与视频处理 · 电气工程与系统科学 2023-06-27 Jiahao Huang , Angelica Aviles-Rivero , Carola-Bibiane Schönlieb , Guang Yang

A low-resolution digital surface model (DSM) features distinctive attributes impacted by noise, sensor limitations and data acquisition conditions, which failed to be replicated using simple interpolation methods like bicubic. This causes…

图像与视频处理 · 电气工程与系统科学 2024-04-08 Daniel Panangian , Ksenia Bittner

Diffusion probabilistic models have been recently used in a variety of tasks, including speech enhancement and synthesis. As a generative approach, diffusion models have been shown to be especially suitable for imputation problems, where…

音频与语音处理 · 电气工程与系统科学 2023-06-05 Tal Peer , Simon Welker , Timo Gerkmann

The pretrained diffusion model as a strong prior has been leveraged to address inverse problems in a zero-shot manner without task-specific retraining. Different from the unconditional generation, the measurement-guided generation requires…

最优化与控制 · 数学 2025-03-14 Ji Li , Chao Wang

The implementation of diffusion-based pansharpening task is predominantly constrained by its slow inference speed, which results from numerous sampling steps. Despite the existing techniques aiming to accelerate sampling, they often…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Shiqi Cao , Liangjian Deng , Shangqi Deng

Score Distillation Sampling (SDS) has emerged as an effective technique for leveraging 2D diffusion priors for tasks such as text-to-3D generation. While powerful, SDS struggles with achieving fine-grained alignment to user intent. To…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Itay Chachy , Guy Yariv , Sagie Benaim

High-quality audio is essential in a wide range of applications, including online communication, virtual assistants, and the multimedia industry. However, degradation caused by noise, compression, and transmission artifacts remains a major…

Recent state-of-the-art image restoration methods mostly adopt latent diffusion models with U-Net backbones, yet still facing challenges in achieving high-quality restoration due to their limited capabilities. Diffusion transformers (DiTs),…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Dehong Kong , Fan Li , Zhixin Wang , Jiaqi Xu , Renjing Pei , Wenbo Li , WenQi Ren

Flow matching is a recent state-of-the-art framework for generative modeling based on ordinary differential equations (ODEs). While closely related to diffusion models, it provides a more general perspective on generative modeling. Although…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Jeongsol Kim , Bryan Sangwoo Kim , Jong Chul Ye

We introduce Audio-SDS, a generalization of Score Distillation Sampling (SDS) to text-conditioned audio diffusion models. While SDS was initially designed for text-to-3D generation using image diffusion, its core idea of distilling a…

声音 · 计算机科学 2025-05-08 Jessie Richter-Powell , Antonio Torralba , Jonathan Lorraine

Anatomically guided PET reconstruction using MRI information has been shown to have the potential to improve PET image quality. However, these improvements are limited to PET scans with paired MRI information. In this work we employed a…

The term "differentiable digital signal processing" describes a family of techniques in which loss function gradients are backpropagated through digital signal processors, facilitating their integration into neural networks. This article…

声音 · 计算机科学 2023-08-30 Ben Hayes , Jordie Shier , György Fazekas , Andrew McPherson , Charalampos Saitis

In recent studies, diffusion models have shown promise as priors for solving audio inverse problems. These models allow us to sample from the posterior distribution of a target signal given an observed signal by manipulating the diffusion…

音频与语音处理 · 电气工程与系统科学 2024-10-22 Chin-Yun Yu , Emilian Postolache , Emanuele Rodolà , György Fazekas

Diffusion models have emerged as a key pillar of foundation models in visual domains. One of their critical applications is to universally solve different downstream inverse tasks via a single diffusion prior without re-training for each…

机器学习 · 计算机科学 2023-10-03 Morteza Mardani , Jiaming Song , Jan Kautz , Arash Vahdat