English
Related papers

Related papers: Source-Filter HiFi-GAN: Fast and Pitch Controllabl…

200 papers

In this work, we propose a new mathematical vocoder algorithm(modified spectral inversion) that generates a waveform from acoustic features without phase estimation. The main benefit of using our proposed method is that it excludes the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-17 Hyun Gon Ryu , Jeong-Hoon Kim , Simon See

Generative adversarial network (GAN) still exists some problems in dealing with speech enhancement (SE) task. Some GAN-based systems adopt the same structure from Pixel-to-Pixel directly without special optimization. The importance of the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-09 Huixiang Huang , Renjie Wu , Jingbiao Huang , Jucai Lin , Jun Yin

Speech Enhancement techniques have become core technologies in mobile devices and voice software. Still, modern deep learning solutions often require high amount of computational resources what makes their usage on low-resource devices…

Sound · Computer Science 2025-07-24 Ekaterina Dmitrieva , Maksim Kaledin

Have you ever wondered how a song might sound if performed by a different artist? In this work, we propose SCM-GAN, an end-to-end non-parallel song conversion system powered by generative adversarial and transfer learning that allows users…

Machine Learning · Computer Science 2020-02-03 Rema Daher , Mohammad Kassem Zein , Julia El Zini , Mariette Awad , Daniel Asmar

Generative Adversarial Networks (GANs) are a powerful class of generative models. Despite their successes, the most appropriate choice of a GAN network architecture is still not well understood. GAN models for image synthesis have adopted a…

Machine Learning · Computer Science 2019-05-28 Sukarna Barua , Sarah Monazam Erfani , James Bailey

This paper proposes Scyclone, a high-quality voice conversion (VC) technique without parallel data training. Scyclone improves speech naturalness and speaker similarity of the converted speech by introducing CycleGAN-based spectrogram…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-08 Masaya Tanaka , Takashi Nose , Aoi Kanagaki , Ryohei Shimizu , Akira Ito

Point clouds acquired from range scans are often sparse, noisy, and non-uniform. This paper presents a new point cloud upsampling network called PU-GAN, which is formulated based on a generative adversarial network (GAN), to learn a rich…

Computer Vision and Pattern Recognition · Computer Science 2019-07-26 Ruihui Li , Xianzhi Li , Chi-Wing Fu , Daniel Cohen-Or , Pheng-Ann Heng

Generative adversarial networks (GANs) have emerged as a powerful paradigm for producing high-fidelity data samples, yet their performance is constrained by the quality of latent representations, typically sampled from classical noise…

Quantum Physics · Physics 2025-08-19 Kun Ming Goh

We present a high-fidelity 3D generative adversarial network (GAN) inversion framework that can synthesize photo-realistic novel views while preserving specific details of the input image. High-fidelity 3D GAN inversion is inherently…

Computer Vision and Pattern Recognition · Computer Science 2022-11-30 Jiaxin Xie , Hao Ouyang , Jingtan Piao , Chenyang Lei , Qifeng Chen

WaveCycleGAN has recently been proposed to bridge the gap between natural and synthesized speech waveforms in statistical parametric speech synthesis and provides fast inference with a moving average model rather than an autoregressive…

Sound · Computer Science 2019-04-10 Kou Tanaka , Hirokazu Kameoka , Takuhiro Kaneko , Nobukatsu Hojo

Deep learning based visual to sound generation systems essentially need to be developed particularly considering the synchronicity aspects of visual and audio features with time. In this research we introduce a novel task of guiding a class…

Machine Learning · Computer Science 2021-07-21 Sanchita Ghose , John J. Prevost

Improving speech system performance in noisy environments remains a challenging task, and speech enhancement (SE) is one of the effective techniques to solve the problem. Motivated by the promising results of generative adversarial networks…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-05 Daniel Michelsanti , Zheng-Hua Tan

Generative adversarial networks (GANs) are challenging to train stably, and a promising remedy of injecting instance noise into the discriminator input has not been very effective in practice. In this paper, we propose Diffusion-GAN, a…

Machine Learning · Computer Science 2023-08-29 Zhendong Wang , Huangjie Zheng , Pengcheng He , Weizhu Chen , Mingyuan Zhou

Pansharpening is a widely used image enhancement technique for remote sensing. Its principle is to fuse the input high-resolution single-channel panchromatic (PAN) image and low-resolution multi-spectral image and to obtain a…

Computer Vision and Pattern Recognition · Computer Science 2021-12-01 Zixiang Zhao , Jiangshe Zhang , Shuang Xu , Kai Sun , Lu Huang , Junmin Liu , Chunxia Zhang

In this paper, we propose a novel controllable text-to-image generative adversarial network (ControlGAN), which can effectively synthesise high-quality images and also control parts of the image generation according to natural language…

Computer Vision and Pattern Recognition · Computer Science 2019-12-20 Bowen Li , Xiaojuan Qi , Thomas Lukasiewicz , Philip H. S. Torr

Flow matching offers a robust and stable approach to training diffusion models. However, directly applying flow matching to neural vocoders can result in subpar audio quality. In this work, we present WaveFM, a reparameterized flow matching…

Sound · Computer Science 2025-03-24 Tianze Luo , Xingchen Miao , Wenbo Duan

Random noise arising from physical processes is an inherent characteristic of measurements and a limiting factor for most signal processing and data analysis tasks. Given the recent interest in generative adversarial networks (GANs) for…

Signal Processing · Electrical Eng. & Systems 2023-08-22 Adam Wunderlich , Jack Sklar

Generative Adversarial Networks (GANs) represent an attractive and novel approach to generate realistic data, such as genes, proteins, or drugs, in synthetic biology. Here, we apply GANs to generate synthetic DNA sequences encoding for…

Genomics · Quantitative Biology 2018-04-06 Anvita Gupta , James Zou

This paper proposes ESTVocoder, a novel excitation-spectral-transformed neural vocoder within the framework of source-filter theory. The ESTVocoder transforms the amplitude and phase spectra of the excitation into the corresponding speech…

Sound · Computer Science 2024-11-19 Xiao-Hang Jiang , Hui-Peng Du , Yang Ai , Ye-Xin Lu , Zhen-Hua Ling

The traditional vocoders have the advantages of high synthesis efficiency, strong interpretability, and speech editability, while the neural vocoders have the advantage of high synthesis quality. To combine the advantages of two vocoders,…

Sound · Computer Science 2022-03-08 Tao Wang , Ruibo Fu , Jiangyan Yi , Jianhua Tao , Zhengqi Wen