中文
相关论文

相关论文: FreGrad: Lightweight and Fast Frequency-aware Diff…

200 篇论文

We propose a new excitation source signal for VOCODERs and an all-pass impulse response for post-processing of synthetic sounds and pre-processing of natural sounds for data-augmentation. The proposed signals are variants of velvet noise,…

Audio super-resolution is a fundamental task that predicts high-frequency components for low-resolution audio, enhancing audio quality in digital applications. Previous methods have limitations such as the limited scope of audio types…

声音 · 计算机科学 2023-09-15 Haohe Liu , Ke Chen , Qiao Tian , Wenwu Wang , Mark D. Plumbley

The diffusion model performs remarkable in generating high-dimensional content but is computationally intensive, especially during training. We propose Progressive Growing of Diffusion Autoencoder (PaGoDA), a novel pipeline that reduces the…

计算机视觉与模式识别 · 计算机科学 2024-10-30 Dongjun Kim , Chieh-Hsin Lai , Wei-Hsiang Liao , Yuhta Takida , Naoki Murata , Toshimitsu Uesaka , Yuki Mitsufuji , Stefano Ermon

Diffusion models have emerged as a formidable tool for training-free conditional generation.However, a key hurdle in inference-time guidance techniques is the need for compute-heavy backpropagation through the diffusion network for…

计算机视觉与模式识别 · 计算机科学 2024-06-05 Nithin Gopalakrishnan Nair , Vishal M Patel

We present a novel high-fidelity real-time neural vocoder called VocGAN. A recently developed GAN-based vocoder, MelGAN, produces speech waveforms in real-time. However, it often produces a waveform that is insufficient in quality or…

音频与语音处理 · 电气工程与系统科学 2020-07-31 Jinhyeok Yang , Junmo Lee , Youngik Kim , Hoonyoung Cho , Injung Kim

Discrete diffusion models have emerged as a powerful class of models and a promising route to fast language generation, but practical implementations typically rely on factored reverse transitions ignoring cross-token dependencies and…

机器学习 · 计算机科学 2026-05-14 Dario Shariatian , Alain Durmus , Umut Simsekli , Stefano Peluchetti

Visual generative models (e.g., diffusion models) typically operate in compressed latent spaces to balance training efficiency and sample quality. In parallel, there has been growing interest in leveraging high-quality pre-trained visual…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Yuan Gao , Chen Chen , Tianrong Chen , Jiatao Gu

Capturing high-level structure in audio waveforms is challenging because a single second of audio spans tens of thousands of timesteps. While long-range dependencies are difficult to model directly in the time domain, we show that they can…

音频与语音处理 · 电气工程与系统科学 2019-06-05 Sean Vasquez , Mike Lewis

Diffusion models are proficient at generating high-quality images. They are however effective only when operating at the resolution used during training. Inference at a scaled resolution leads to repetitive patterns and structural…

计算机视觉与模式识别 · 计算机科学 2024-11-28 Haosen Yang , Adrian Bulat , Isma Hadji , Hai X. Pham , Xiatian Zhu , Georgios Tzimiropoulos , Brais Martinez

Generative adversarial network (GAN) based vocoders have achieved significant attention in speech synthesis with high quality and fast inference speed. However, there still exist many noticeable spectral artifacts, resulting in the quality…

音频与语音处理 · 电气工程与系统科学 2024-07-08 Rubing Shen , Yanzhen Ren , Zongkun Sun

Generating high-fidelity visual effects (VFX) typically demands massive datasets and prohibitive computational power due to the intricate coupling of spatial textures and temporal dynamics. In this paper, we introduce EasyVFX, a…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Yue Ma , Xu Ye , Qinghe Wang , Yucheng Wang , Hongyu Liu , Yinhan Zhang , Xinyu Wang , Yuanpeng Che , Shanhui Mo , Paul Liang , Fangneng Zhan , Qifeng Chen

The introduction of diffusion models has brought significant advances to the field of audio-driven talking head generation. However, the extremely slow inference speed severely limits the practical implementation of diffusion-based talking…

图形学 · 计算机科学 2025-11-18 Haotian Wang , Yuzhe Weng , Jun Du , Haoran Xu , Xiaoyan Wu , Shan He , Bing Yin , Cong Liu , Jianqing Gao , Qingfeng Liu

Distilled diffusion models accelerate image generation by reducing the number of denoising steps, but often suffer from degraded image quality. To mitigate this trade-off, test-time optimization methods improve quality, yet their iterative…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Haewon Jeon , Si-Hyeon Lee

Text-driven video generation has advanced significantly due to developments in diffusion models. Beyond the training and sampling phases, recent studies have investigated noise priors of diffusion models, as improved noise priors yield…

图像与视频处理 · 电气工程与系统科学 2025-02-20 Yunlong Yuan , Yuanfan Guo , Chunwei Wang , Wei Zhang , Hang Xu , Li Zhang

Recently, GAN-based neural vocoders such as Parallel WaveGAN, MelGAN, HiFiGAN, and UnivNet have become popular due to their lightweight and parallel structure, resulting in a real-time synthesized waveform with high fidelity, even on a CPU.…

声音 · 计算机科学 2022-06-22 Yi Wang , Yi Si

Expressive text-to-speech systems have undergone significant advancements owing to prosody modeling, but conventional methods can still be improved. Traditional approaches have relied on the autoregressive method to predict the quantized…

声音 · 计算机科学 2025-01-22 Hyung-Seok Oh , Sang-Hoon Lee , Seong-Whan Lee

This paper addresses the challenge of enhancing the realism of vocoder-generated singing voice audio by mitigating the distinguishable disparities between synthetic and real-life recordings, particularly in high-frequency spectrogram…

声音 · 计算机科学 2025-08-05 Runxuan Yang , Kai Li , Guo Chen , Xiaolin Hu

Diffusion speech enhancement on discrete audio codec features gain immense attention due to their improved speech component reconstruction capability. However, they usually suffer from high inference computational complexity due to multiple…

音频与语音处理 · 电气工程与系统科学 2026-01-30 Yihui Fu , Tim Fingscheidt

Speech tokenizers serve as foundational components for speech language models, yet current designs exhibit several limitations, including: 1) dependence on multi-layer residual vector quantization structures or high frame rates, 2) reliance…

声音 · 计算机科学 2025-08-26 Yuancheng Wang , Dekun Chen , Xueyao Zhang , Junan Zhang , Jiaqi Li , Zhizheng Wu

Unsupervised representation learning of speech has been of keen interest in recent years, which is for example evident in the wide interest of the ZeroSpeech challenges. This work presents a new method for learning frame level…

音频与语音处理 · 电气工程与系统科学 2020-08-18 Mingjie Chen , Thomas Hain