English
Related papers

Related papers: PostGAN: A GAN-Based Post-Processor to Enhance the…

200 papers

Non-autoregressive GAN-based neural vocoders are widely used due to their fast inference speed and high perceptual quality. However, they often suffer from audible artifacts such as tonal artifacts in their generated results. Therefore, we…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-11 Hyunjae Cho , Junhyeok Lee , Wonbin Jung

Our previous work, the unified source-filter GAN (uSFGAN) vocoder, introduced a novel architecture based on the source-filter theory into the parallel waveform generative adversarial network to achieve high voice quality and pitch…

Sound · Computer Science 2023-02-28 Reo Yoneyama , Yi-Chiao Wu , Tomoki Toda

Recent studies have shown that neural vocoders based on generative adversarial network (GAN) can generate audios with high quality. While GAN based neural vocoders have shown to be computationally much more efficient than those based on…

Sound · Computer Science 2021-06-28 Zhengxi Liu , Yanmin Qian

While existing speech audio codecs designed for compression exploit limited forms of temporal redundancy and allow for multi-scale representations, they tend to represent all features of audio in the same way. In contrast, generative voice…

Sound · Computer Science 2025-09-22 Ryan Collette , Ross Greenwood , Serena Nicoll

Real-world audio recordings are often degraded by factors such as noise, reverberation, and equalization distortion. This paper introduces HiFi-GAN, a deep learning method to transform recorded speech to sound as though it had been recorded…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-23 Jiaqi Su , Zeyu Jin , Adam Finkelstein

In this work, we propose a full-band real-time speech enhancement system with GAN-based stochastic regeneration. Predictive models focus on estimating the mean of the target distribution, whereas generative models aim to learn the full…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-30 Sanberk Serbest , Tijana Stojkovic , Milos Cernak , Andrew Harper

In this paper, we propose multi-band MelGAN, a much faster waveform generation model targeting to high-quality text-to-speech. Specifically, we improve the original MelGAN by the following aspects. First, we increase the receptive field of…

Sound · Computer Science 2020-11-18 Geng Yang , Shan Yang , Kai Liu , Peng Fang , Wei Chen , Lei Xie

This paper presents a speaking-rate-controllable HiFi-GAN neural vocoder. Original HiFi-GAN is a high-fidelity, computationally efficient, and tiny-footprint neural vocoder. We attempt to incorporate a speaking rate control function into…

Sound · Computer Science 2022-04-25 Detai Xin , Shinnosuke Takamichi , Takuma Okamoto , Hisashi Kawai , Hiroshi Saruwatari

Enhancing speech quality under adverse SNR conditions remains a significant challenge for discriminative deep neural network (DNN)-based approaches. In this work, we propose DisCoGAN, which is a time-frequency-domain generative adversarial…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-18 Shrishti Saha Shetu , Emanuël A. P. Habets , Andreas Brendel

Speech enhancement at extremely low signal-to-noise ratio (SNR) condition is a very challenging problem and rarely investigated in previous works. This paper proposes a robust speech enhancement approach (UNetGAN) based on U-Net and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-30 Xiang Hao , Xiangdong Su , Zhiyu Wang , Hui Zhang , Batushiren

Generative Adversarial Networks (GANs) have become exceedingly popular in a wide range of data-driven research fields, due in part to their success in image generation. Their ability to generate new samples, often from only a small amount…

Computation and Language · Computer Science 2019-03-19 Thomas Wiest , Nicholas Cummins , Alice Baird , Simone Hantke , Judith Dineley , Björn Schuller

Since the introduction of Generative Adversarial Networks (GANs) in speech synthesis, remarkable achievements have been attained. In a thorough exploration of vocoders, it has been discovered that audio waveforms can be generated at speeds…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-14 Yubing Cao , Yongming Li , Liejun Wang , Yinfeng Yu

Several recent work on speech synthesis have employed generative adversarial networks (GANs) to produce raw waveforms. Although such methods improve the sampling efficiency and memory usage, their sample quality has not yet reached that of…

Sound · Computer Science 2020-10-26 Jungil Kong , Jaehyeon Kim , Jaekyoung Bae

Packet loss is a major cause of voice quality degradation in VoIP transmissions with serious impact on intelligibility and user experience. This paper describes a system based on a generative adversarial approach, which aims to repair the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-31 Carlo Aironi , Samuele Cornell , Luca Serafini , Stefano Squartini

Improving speech system performance in noisy environments remains a challenging task, and speech enhancement (SE) is one of the effective techniques to solve the problem. Motivated by the promising results of generative adversarial networks…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-05 Daniel Michelsanti , Zheng-Hua Tan

The generative adversarial networks (GANs) have facilitated the development of speech enhancement recently. Nevertheless, the performance advantage is still limited when compared with state-of-the-art models. In this paper, we propose a…

Sound · Computer Science 2020-06-16 Andong Li , Chengshi Zheng , Renhua Peng , Cunhang Fan , Xiaodong Li

Recently, BigVGAN has emerged as high-performance speech vocoder. Its sequence-to-sequence-based synthesis, however, prohibits usage in low-latency conversational applications. Our work addresses this shortcoming in three steps. First, we…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-27 Renzheng Shi , Andreas Bär , Marvin Sach , Wouter Tirry , Tim Fingscheidt

We propose a unified signal compression framework that uses a generative adversarial network (GAN) to compress heterogeneous signals. The compressed signal is represented as a latent vector and fed into a generator network that is trained…

Signal Processing · Electrical Eng. & Systems 2021-09-24 Bowen Liu , Changwoo Lee , Ang Cao , Hun-Seok Kim

Recently, GAN based speech synthesis methods, such as MelGAN, have become very popular. Compared to conventional autoregressive based methods, parallel structures based generators make waveform generation process fast and stable. However,…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-25 Qiao Tian , Yi Chen , Zewang Zhang , Heng Lu , Linghui Chen , Lei Xie , Shan Liu

We present a neural vocoder designed with low-powered Alternative and Augmentative Communication devices in mind. By combining elements of successful modern vocoders with established ideas from an older generation of technology, our system…

Sound · Computer Science 2023-06-09 Oliver Watts , Lovisa Wihlborg , Cassia Valentini-Botinhao