中文
相关论文

相关论文: A2SB: Audio-to-Audio Schrodinger Bridges

200 篇论文

We propose Image-to-Image Schr\"odinger Bridge (I$^2$SB), a new class of conditional diffusion models that directly learn the nonlinear diffusion processes between two given distributions. These diffusion bridges are particularly useful for…

计算机视觉与模式识别 · 计算机科学 2023-05-29 Guan-Horng Liu , Arash Vahdat , De-An Huang , Evangelos A. Theodorou , Weili Nie , Anima Anandkumar

This paper proposes a generative speech enhancement model based on Schr\"odinger bridge (SB). The proposed model is employing a tractable SB to formulate a data-to-data process between the clean speech distribution and the observed noisy…

音频与语音处理 · 电气工程与系统科学 2024-07-24 Ante Jukić , Roman Korostik , Jagadeesh Balam , Boris Ginsburg

Speech super-resolution (SR), which generates a waveform at a higher sampling rate from its low-resolution version, is a long-standing critical task in speech restoration. Previous works have explored speech SR in different data spaces, but…

声音 · 计算机科学 2025-01-15 Chang Li , Zehua Chen , Fan Bao , Jun Zhu

In this work, we investigate application of generative speech enhancement to improve the robustness of ASR models in noisy and reverberant conditions. We employ a recently-proposed speech enhancement model based on Schr\"odinger bridge,…

音频与语音处理 · 电气工程与系统科学 2025-05-09 Rauf Nasretdinov , Roman Korostik , Ante Jukić

Recently, diffusion-based generative models have demonstrated remarkable performance in speech enhancement tasks. However, these methods still encounter challenges, including the lack of structural information and poor performance in low…

音频与语音处理 · 电气工程与系统科学 2024-09-16 Siyi Wang , Siyi Liu , Andrew Harper , Paul Kendrick , Mathieu Salzmann , Milos Cernak

With active research in audio compression techniques yielding substantial breakthroughs, spectral reconstruction of low-quality audio waves remains a less indulged topic. In this paper, we propose a novel approach for reconstructing higher…

声音 · 计算机科学 2021-08-10 Darshan Deshpande , Harshavardhan Abichandani

Image inpainting is an important image generation task, which aims to restore corrupted image from partial visible area. Recently, diffusion Schr\"odinger bridge methods effectively tackle this task by modeling the translation between…

计算机视觉与模式识别 · 计算机科学 2024-12-12 Zihao Han , Baoquan Zhang , Lisai Zhang , Shanshan Feng , Kenghong Lin , Guotao Liang , Yunming Ye , Xiaochen Qi , Guangming Ye

Diffusion models serve as a powerful generative framework for solving inverse problems. However, they still face two key challenges: 1) the distortion-perception tradeoff, where improving perceptual quality often degrades reconstruction…

机器学习 · 计算机科学 2025-11-20 Qing Yao , Lijian Gao , Qirong Mao , Ming Dong

Diffusion-based models have demonstrated remarkable effectiveness in image restoration tasks; however, their iterative denoising process, which starts from Gaussian noise, often leads to slow inference speeds. The Image-to-Image…

图像与视频处理 · 电气工程与系统科学 2025-03-25 Yuang Wang , Siyeop Yoon , Pengfei Jin , Matthew Tivnan , Sifan Song , Zhennong Chen , Rui Hu , Li Zhang , Quanzheng Li , Zhiqiang Chen , Dufan Wu

Audio super-resolution is a fundamental task that predicts high-frequency components for low-resolution audio, enhancing audio quality in digital applications. Previous methods have limitations such as the limited scope of audio types…

声音 · 计算机科学 2023-09-15 Haohe Liu , Ke Chen , Qiao Tian , Wenwu Wang , Mark D. Plumbley

Audio inpainting aims to reconstruct missing segments in corrupted recordings. Most of existing methods produce plausible reconstructions when the gap lengths are short, but struggle to reconstruct gaps larger than about 100 ms. This paper…

音频与语音处理 · 电气工程与系统科学 2025-01-13 Eloi Moliner , Vesa Välimäki

Audio inpainting seeks to restore missing segments in degraded recordings. Previous diffusion-based methods exhibit impaired performance when the missing region is large. We introduce the first approach that applies discrete diffusion over…

声音 · 计算机科学 2026-02-18 Tali Dror , Iftach Shoham , Moshe Buchris , Oren Gal , Haim Permuter , Gilad Katz , Eliya Nachmani

Diffusion models are a powerful class of generative models which simulate stochastic differential equations (SDEs) to generate data from noise. While diffusion models have achieved remarkable progress, they have limitations in unpaired…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Beomsu Kim , Gihyun Kwon , Kwanyoung Kim , Jong Chul Ye

Considering the microphone is easily affected by noise and soundproof materials, the radio frequency (RF) signal is a promising candidate to recover audio as it is immune to noise and can traverse many soundproof objects. In this paper, we…

声音 · 计算机科学 2022-06-23 Running Zhao , Jiangtao Yu , Tingle Li , Hang Zhao , Edith C. H. Ngai

Long (> 200 ms) audio inpainting, to recover a long missing part in an audio segment, could be widely applied to audio editing tasks and transmission loss recovery. It is a very challenging problem due to the high dimensional, complex and…

声音 · 计算机科学 2019-11-18 Ya-Liang Chang , Kuan-Ying Lee , Po-Yu Wu , Hung-yi Lee , Winston Hsu

Audio-guided face reenactment aims at generating photorealistic faces using audio information while maintaining the same facial movement as when speaking to a real person. However, existing methods can not generate vivid face images or only…

计算机视觉与模式识别 · 计算机科学 2020-05-01 Jiangning Zhang , Liang Liu , Zhucun Xue , Yong Liu

Computed tomography (CT) is a cornerstone imaging modality for non-invasive, high-resolution visualization of internal anatomical structures. However, when the scanned object exceeds the scanner's field of view (FOV), projection data are…

图像与视频处理 · 电气工程与系统科学 2026-04-09 Zhenhao Li , Song Ni , Long Yang , Xiaojie Yin , Haijun Yu , Jiazhou Wang , Hongbin Han , Weigang Hu , Yixing Huang

Piano audio-to-score transcription (A2S) is an important yet underexplored task with extensive applications for music composition, practice, and analysis. However, existing end-to-end piano A2S systems faced difficulties in retrieving…

声音 · 计算机科学 2024-05-24 Wei Zeng , Xian He , Ye Wang

Deep generative models have recently been employed for speech enhancement to generate perceptually valid clean speech on large-scale datasets. Several diffusion models have been proposed, and more recently, a tractable Schr\"odinger Bridge…

声音 · 计算机科学 2025-06-03 Seungu Han , Sungho Lee , Juheon Lee , Kyogu Lee

Magnetic Resonance Imaging (MRI) is an inherently multi-contrast modality, where cross-contrast priors can be exploited to improve image reconstruction from undersampled data. Recently, diffusion models have shown remarkable performance in…

图像与视频处理 · 电气工程与系统科学 2025-10-27 Yue Wang , Yuanbiao Yang , Zhuo-xu Cui , Tian Zhou , Bingsheng Huang , Hairong Zheng , Dong Liang , Yanjie Zhu
‹ 上一页 1 2 3 10 下一页 ›