English
Related papers

Related papers: Parallel WaveGAN: A fast waveform generation model…

200 papers

Speech enhancement aims to obtain speech signals with high intelligibility and quality from noisy speech. Recent work has demonstrated the excellent performance of time-domain deep learning methods, such as Conv-TasNet. However, these…

Sound · Computer Science 2021-09-21 Feiyang Xiao , Jian Guan , Qiuqiang Kong , Wenwu Wang

Building a voice conversion (VC) system from non-parallel speech corpora is challenging but highly valuable in real application scenarios. In most situations, the source and the target speakers do not repeat the same texts or they may even…

Computation and Language · Computer Science 2017-06-09 Chin-Cheng Hsu , Hsin-Te Hwang , Yi-Chiao Wu , Yu Tsao , Hsin-Min Wang

Modern audio generation predominantly relies on latent-space compression, introducing additional complexity and potential information loss. In this work, we challenge this paradigm with WavFlow, a framework that generates high-fidelity…

Generative adversarial networks (GAN) have shown remarkable results in image generation tasks. High fidelity class-conditional GAN methods often rely on stabilization techniques by constraining the global Lipschitz continuity. Such…

Machine Learning · Computer Science 2020-08-11 Jiachen Zhong , Xuanqing Liu , Cho-Jui Hsieh

Learning disentangled representation of data without supervision is an important step towards improving the interpretability of generative models. Despite recent advances in disentangled representation learning, existing approaches often…

Computer Vision and Pattern Recognition · Computer Science 2020-01-14 Wonkwang Lee , Donggyun Kim , Seunghoon Hong , Honglak Lee

We study the ability of Wasserstein Generative Adversarial Network (WGAN) to generate missing audio content which is, in context, (statistically similar) to the sound and the neighboring borders. We deal with the challenge of audio…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-18 P. P. Ebner , A. Eltelt

Neural vocoders based on the generative adversarial neural network (GAN) have been widely used due to their fast inference speed and lightweight networks while generating high-quality speech waveforms. Since the perceptually important…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-04 Taejun Bak , Junmo Lee , Hanbin Bae , Jinhyeok Yang , Jae-Sung Bae , Young-Sun Joo

The diffusion model is capable of generating high-quality data through a probabilistic approach. However, it suffers from the drawback of slow generation speed due to the requirement of a large number of time steps. To address this…

Sound · Computer Science 2024-04-30 Myeongjin Ko , Yong-Hoon Choi

The past few years have witnessed fast development in video quality enhancement via deep learning. Existing methods mainly focus on enhancing the objective quality of compressed video while ignoring its perceptual quality. In this paper, we…

Image and Video Processing · Electrical Eng. & Systems 2020-08-04 Jianyi Wang , Xin Deng , Mai Xu , Congyong Chen , Yuhang Song

This paper presents the development of a generative adversarial network (GAN) for synthesizing dental panoramic radiographs. Although exploratory in nature, the study aims to address the scarcity of data in dental research and education. We…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Soren Pedersen , Sanyam Jain , Mikkel Chavez , Viktor Ladehoff , Bruna Neves de Freitas , Ruben Pauwels

In recent years, generative adversarial networks (GANs) have made significant progress in generating audio sequences. However, these models typically rely on bandwidth-limited mel-spectrograms, which constrain the resolution of generated…

Sound · Computer Science 2025-05-15 Zeeshan Ahmad , Shudi Bao , Meng Chen

Nowadays vast amounts of speech data are recorded from low-quality recorder devices such as smartphones, tablets, laptops, and medium-quality microphones. The objective of this research was to study the automatic generation of high-quality…

Whispered speech is a special way of pronunciation without using vocal cord vibration. A whispered speech does not contain a fundamental frequency, and its energy is about 20dB lower than that of a normal speech. Converting a whispered…

Sound · Computer Science 2021-11-03 Teng Gao , Jian Zhou , Huabin Wang , Liang Tao , Hon Keung Kwan

Improving speech system performance in noisy environments remains a challenging task, and speech enhancement (SE) is one of the effective techniques to solve the problem. Motivated by the promising results of generative adversarial networks…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-05 Daniel Michelsanti , Zheng-Hua Tan

The advancement of high-performance computing has enabled the generation of large direct numerical simulation (DNS) datasets of turbulent flows, driving the need for efficient compression/decompression techniques that reduce storage demands…

We investigate the effectiveness of generative adversarial networks (GANs) for speech enhancement, in the context of improving noise robustness of automatic speech recognition (ASR) systems. Prior work demonstrates that GANs can effectively…

Sound · Computer Science 2018-11-01 Chris Donahue , Bo Li , Rohit Prabhavalkar

Adversarial attacks from generative models often produce low-quality images and require substantial computational resources. Diffusion models, though capable of high-quality generation, typically need hundreds of sampling steps for…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Susim Roy , Anubhooti Jain , Mayank Vatsa , Richa Singh

Recently, universal waveform generation tasks have been investigated conditioned on various out-of-distribution scenarios. Although GAN-based methods have shown their strength in fast waveform generation, they are vulnerable to…

Sound · Computer Science 2024-08-15 Sang-Hoon Lee , Ha-Yeong Choi , Seong-Whan Lee

Adversarial waveform generation has been a popular approach as the backend of singing voice conversion (SVC) to generate high-quality singing audio. However, the instability of GAN also leads to other problems, such as pitch jitters and U/V…

Sound · Computer Science 2022-01-26 Haohan Guo , Zhiping Zhou , Fanbo Meng , Kai Liu

We propose an application of sequence generative adversarial networks (SeqGAN), which are generative adversarial networks for discrete sequence generation, for creating polyphonic musical sequences. Instead of a monophonic melody generation…

Sound · Computer Science 2018-07-03 Sang-gil Lee , Uiwon Hwang , Seonwoo Min , Sungroh Yoon
‹ Prev 1 3 4 5 6 7 10 Next ›