English
Related papers

Related papers: Is GAN Necessary for Mel-Spectrogram-based Neural …

200 papers

Adversarial loss in a conditional generative adversarial network (GAN) is not designed to directly optimize evaluation metrics of a target task, and thus, may not always guide the generator in a GAN to generate data with improved metric…

Sound · Computer Science 2019-05-14 Szu-Wei Fu , Chien-Feng Liao , Yu Tsao , Shou-De Lin

One way to interpret trained deep neural networks (DNNs) is by inspecting characteristics that neurons in the model respond to, such as by iteratively optimising the model input (e.g., an image) to maximally activate specific neurons.…

Machine Learning · Computer Science 2019-07-02 Saumitra Mishra , Daniel Stoller , Emmanouil Benetos , Bob L. Sturm , Simon Dixon

This work adapts two recent architectures of generative models and evaluates their effectiveness for the conversion of whispered speech to normal speech. We incorporate the normal target speech into the training criterion of…

Generative adversarial networks (GANs) are one of the most widely used generative models. GANs can learn complex multi-modal distributions, and generate real-like samples. Despite the major success of GANs in generating synthetic data, they…

Machine Learning · Computer Science 2021-09-07 Sanaz Mohammadjafari , Mucahit Cevik , Ayse Basar

Traditional voice conversion methods rely on parallel recordings of multiple speakers pronouncing the same sentences. For real-world applications however, parallel data is rarely available. We propose MelGAN-VC, a voice conversion method…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-06 Marco Pasini

Voice conversion is a method that allows for the transformation of speaking style while maintaining the integrity of linguistic information. There are many researchers using deep generative models for voice conversion tasks. Generative…

Sound · Computer Science 2023-08-29 Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao

High-fidelity singing voices usually require higher sampling rate (e.g., 48kHz) to convey expression and emotion. However, higher sampling rate causes the wider frequency band and longer waveform sequences and throws challenges for singing…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-04 Jiawei Chen , Xu Tan , Jian Luan , Tao Qin , Tie-Yan Liu

Encoder-decoder GANs architectures (e.g., BiGAN and ALI) seek to add an inference mechanism to the GANs setup, consisting of a small encoder deep net that maps data-points to their succinct encodings. The intuition is that being forced to…

Machine Learning · Computer Science 2017-11-08 Sanjeev Arora , Andrej Risteski , Yi Zhang

Current speech enhancement techniques operate on the spectral domain and/or exploit some higher-level feature. The majority of them tackle a limited number of noise conditions and rely on first-order statistics. To circumvent these issues,…

Machine Learning · Computer Science 2017-06-12 Santiago Pascual , Antonio Bonafonte , Joan Serrà

This paper proposes a spectral-domain perceptual weighting technique for Parallel WaveGAN-based text-to-speech (TTS) systems. The recently proposed Parallel WaveGAN vocoder successfully generates waveform sequences using a fast…

Audio and Speech Processing · Electrical Eng. & Systems 2021-01-20 Eunwoo Song , Ryuichi Yamamoto , Min-Jae Hwang , Jin-Seob Kim , Ohsung Kwon , Jae-Min Kim

Auto-encoding generative adversarial networks (GANs) combine the standard GAN algorithm, which discriminates between real and model-generated data, with a reconstruction loss given by an auto-encoder. Such models aim to prevent mode…

Machine Learning · Statistics 2017-10-24 Mihaela Rosca , Balaji Lakshminarayanan , David Warde-Farley , Shakir Mohamed

Generative Adversarial Networks (GAN) have received wide attention in the machine learning field for their potential to learn high-dimensional, complex real data distribution. Specifically, they do not rely on any assumptions about the…

Machine Learning · Computer Science 2019-03-01 Yongjun Hong , Uiwon Hwang , Jaeyoon Yoo , Sungroh Yoon

High-quality recordings of radio frequency (RF) emissions from commercial communication hardware in realistic environments are often needed to develop and assess spectrum-sharing technologies and practices, e.g., for training and testing…

Signal Processing · Electrical Eng. & Systems 2022-02-21 Jack Sklar , Adam Wunderlich

Standard neural networks are often overconfident when presented with data outside the training distribution. We introduce HyperGAN, a new generative model for learning a distribution of neural network parameters. HyperGAN does not require…

Machine Learning · Computer Science 2020-07-16 Neale Ratzlaff , Li Fuxin

Deep learning has become a standard approach for the modeling of audio effects, yet strictly black-box modeling remains problematic for time-varying systems. Unlike time-invariant effects, training models on devices with internal modulation…

Sound · Computer Science 2025-12-18 Yann Bourdin , Pierrick Legrand , Fanny Roche

A generative adversarial network (GAN)-based vocoder trained with an adversarial discriminator is commonly used for speech synthesis because of its fast, lightweight, and high-quality characteristics. However, this data-driven model…

Sound · Computer Science 2024-03-26 Takuhiro Kaneko , Hirokazu Kameoka , Kou Tanaka

Recent advancements in speech synthesis have leveraged GAN-based networks like HiFi-GAN and BigVGAN to produce high-fidelity waveforms from mel-spectrograms. However, these networks are computationally expensive and parameter-heavy.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-19 Yinghao Aaron Li , Cong Han , Xilin Jiang , Nima Mesgarani

Generative Adversarial Networks (GANs) have the ability to generate images that are visually indistinguishable from real images. However, recent studies have revealed that generated and real images share significant differences in the…

Computer Vision and Pattern Recognition · Computer Science 2021-12-01 Ziqiang Li , Pengfei Xia , Xue Rui , Bin Li

Various deepfake detectors have been proposed, but challenges still exist to detect images of unknown categories or GAN models outside of the training settings. Such issues arise from the overfitting issue, which we discover from our own…

Computer Vision and Pattern Recognition · Computer Science 2022-02-08 Yonghyun Jeong , Doyeon Kim , Youngmin Ro , Jongwon Choi

Speaking rate refers to the average number of phonemes within some unit time, while the rhythmic patterns refer to duration distributions for realizations of different phonemes within different phonetic structures. Both are key components…

Sound · Computer Science 2018-08-10 Cheng-chieh Yeh , Po-chun Hsu , Ju-chieh Chou , Hung-yi Lee , Lin-shan Lee