中文
相关论文

相关论文: WSRGlow: A Glow-based Waveform Generative Model fo…

200 篇论文

Implicit neural representations (INRs) are a rapidly growing research field, which provides alternative ways to represent multimedia signals. Recent applications of INRs include image super-resolution, compression of high-dimensional…

Super-resolution (SR) is an ill-posed inverse problem, where the size of the set of feasible solutions that are consistent with a given low-resolution image is very large. Many algorithms have been proposed to find a "good" solution among…

图像与视频处理 · 电气工程与系统科学 2024-03-01 Cansu Korkmaz , A. Murat Tekalp , Zafer Dogan

State-of-the-art under-determined audio source separation systems rely on supervised end-end training of carefully tailored neural network architectures operating either in the time or the spectral domain. However, these methods are…

音频与语音处理 · 电气工程与系统科学 2020-05-29 Vivek Narayanaswamy , Jayaraman J. Thiagarajan , Rushil Anirudh , Andreas Spanias

This paper proposes a generative pretraining foundation model for high-quality speech restoration tasks. By directly operating on complex-valued short-time Fourier transform coefficients, our model does not rely on any vocoders for…

音频与语音处理 · 电气工程与系统科学 2024-09-26 Pin-Jui Ku , Alexander H. Liu , Roman Korostik , Sung-Feng Huang , Szu-Wei Fu , Ante Jukić

We propose WaveTrainerFit, a neural vocoder that performs high-quality waveform generation from data-driven features such as SSL features. WaveTrainerFit builds upon the WaveFit vocoder, which integrates diffusion model and generative…

音频与语音处理 · 电气工程与系统科学 2026-02-06 Hien Ohnaka , Yuma Shirahata , Masaya Kawamura

Versatile audio super-resolution (SR) aims to predict high-frequency components from low-resolution audio across diverse domains such as speech, music, and sound effects. Existing diffusion-based SR methods often fail to produce…

音频与语音处理 · 电气工程与系统科学 2025-09-30 Jaekwon Im , Juhan Nam

Most modern text-to-speech architectures use a WaveNet vocoder for synthesizing high-fidelity waveform audio, but there have been limitations, such as high inference time, in its practical application due to its ancestral sampling scheme.…

声音 · 计算机科学 2019-05-21 Sungwon Kim , Sang-gil Lee , Jongyoon Song , Jaehyeon Kim , Sungroh Yoon

Sequential models achieve state-of-the-art results in audio, visual and textual domains with respect to both estimating the data distribution and generating high-quality samples. Efficient sampling for this class of models has however…

For image super-resolution (SR), bridging the gap between the performance on synthetic datasets and real-world degradation scenarios remains a challenge. This work introduces a novel "Low-Res Leads the Way" (LWay) training framework,…

图像与视频处理 · 电气工程与系统科学 2024-03-06 Haoyu Chen , Wenbo Li , Jinjin Gu , Jingjing Ren , Haoze Sun , Xueyi Zou , Zhensong Zhang , Youliang Yan , Lei Zhu

Speech super-resolution (SR), which generates a waveform at a higher sampling rate from its low-resolution version, is a long-standing critical task in speech restoration. Previous works have explored speech SR in different data spaces, but…

声音 · 计算机科学 2025-01-15 Chang Li , Zehua Chen , Fan Bao , Jun Zhu

Wireless channel modeling plays a pivotal role in designing, analyzing, and optimizing wireless communication systems. Nevertheless, developing an effective channel modeling approach has been a long-standing challenge. This issue has been…

网络与互联网体系结构 · 计算机科学 2025-03-25 Chaozheng Wen , Jingwen Tong , Yingdong Hu , Zehong Lin , Jun Zhang

This paper addresses the standard generalized likelihood ratio test (GLRT) detection problem of weak signals in background noise. In so doing, we consider a nonfluctuating target embedded in complex white Gaussian noise (CWGN), in which the…

Recently, autoregressive neural vocoders have provided remarkable performance in generating high-fidelity speech and have been able to produce synthetic speech in real-time. However, autoregressive neural vocoders such as WaveFlow are…

声音 · 计算机科学 2022-03-28 Manh Luong , Viet Anh Tran

Traditional structured prediction models try to learn the conditional likelihood, i.e., p(y|x), to capture the relationship between the structured output y and the input features x. For many models, computing the likelihood is intractable.…

机器学习 · 计算机科学 2020-02-28 You Lu , Bert Huang

Most learning-based super-resolution (SR) methods aim to recover high-resolution (HR) image from a given low-resolution (LR) image via learning on LR-HR image pairs. The SR methods learned on synthetic data do not perform well in…

图像与视频处理 · 电气工程与系统科学 2020-01-09 Dong Gong , Wei Sun , Qinfeng Shi , Anton van den Hengel , Yanning Zhang

High-resolution image synthesis remains a core challenge in generative modeling, particularly in balancing computational efficiency with the preservation of fine-grained visual detail. We present Latent Wavelet Diffusion (LWD), a…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Luigi Sigillo , Shengfeng He , Danilo Comminiello

MRI super-resolution (SR) and denoising tasks are fundamental challenges in the field of deep learning, which have traditionally been treated as distinct tasks with separate paired training data. In this paper, we propose an innovative…

图像与视频处理 · 电气工程与系统科学 2023-08-24 Qi Wang , Lucas Mahler , Julius Steiglechner , Florian Birk , Klaus Scheffler , Gabriele Lohmann

Learning based single image super-resolution (SISR) for real-world images has been an active research topic yet a challenging task, due to the lack of paired low-resolution (LR) and high-resolution (HR) training images. Most of the existing…

计算机视觉与模式识别 · 计算机科学 2024-09-30 Wanjie Sun , Zhenzhong Chen

Amid the burgeoning development of generative models like diffusion models, the task of differentiating synthesized audio from its natural counterpart grows more daunting. Deepfake detection offers a viable solution to combat this…

密码学与安全 · 计算机科学 2024-07-18 Weizhi Liu , Yue Li , Dongdong Lin , Hui Tian , Haizhou Li

Neural waveform models such as the WaveNet are used in many recent text-to-speech systems, but the original WaveNet is quite slow in waveform generation because of its autoregressive (AR) structure. Although faster non-AR models were…

音频与语音处理 · 电气工程与系统科学 2019-04-30 Xin Wang , Shinji Takaki , Junichi Yamagishi