English
Related papers

Related papers: Unified Source-Filter GAN with Harmonic-plus-Noise…

200 papers

Recently, mainstream mel-spectrogram-based neural vocoders rely on generative adversarial network (GAN) for high-fidelity speech generation, e.g., HiFi-GAN and BigVGAN. However, the use of GAN restricts training efficiency and model…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-12 Hui-Peng Du , Yang Ai , Rui-Chen Zheng , Ye-Xin Lu , Zhen-Hua Ling

Recently, GAN-based neural vocoders such as Parallel WaveGAN, MelGAN, HiFiGAN, and UnivNet have become popular due to their lightweight and parallel structure, resulting in a real-time synthesized waveform with high fidelity, even on a CPU.…

Sound · Computer Science 2022-06-22 Yi Wang , Yi Si

Generative adversarial network (GAN)-based vocoders have been intensively studied because they can synthesize high-fidelity audio waveforms faster than real-time. However, it has been reported that most GANs fail to obtain the optimal…

Sound · Computer Science 2024-03-26 Takashi Shibuya , Yuhta Takida , Yuki Mitsufuji

Automatic speech recognition (ASR) systems are of vital importance nowadays in commonplace tasks such as speech-to-text processing and language translation. This created the need for an ASR system that can operate in realistic crowded…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-29 Sherif Abdulatif , Karim Armanious , Karim Guirguis , Jayasankar T. Sajeev , Bin Yang

Building a voice conversion system for noisy target speakers, such as users providing noisy samples or Internet found data, is a challenging task since the use of contaminated speech in model training will apparently degrade the conversion…

Sound · Computer Science 2022-07-05 Liumeng Xue , Shan Yang , Na Hu , Dan Su , Lei Xie

Universal speech enhancement aims at handling inputs with various speech distortions and recording conditions. In this work, we propose a novel hybrid architecture that synergizes the signal fidelity of discriminative modeling with the…

Sound · Computer Science 2026-01-28 Yinghao Liu , Chengwei Liu , Xiaotao Liang , Haoyin Yan , Shaofei Xue , Zheng Xue

Conditional generative adversarial networks (cGAN) have led to large improvements in the task of conditional image generation, which lies at the heart of computer vision. The major focus so far has been on performance improvement, while…

Machine Learning · Computer Science 2019-03-14 Grigorios G. Chrysos , Jean Kossaifi , Stefanos Zafeiriou

The speech enhancement task usually consists of removing additive noise or reverberation that partially mask spoken utterances, affecting their intelligibility. However, little attention is drawn to other, perhaps more aggressive signal…

Sound · Computer Science 2019-04-09 Santiago Pascual , Joan Serrà , Antonio Bonafonte

Current speech enhancement techniques operate on the spectral domain and/or exploit some higher-level feature. The majority of them tackle a limited number of noise conditions and rely on first-order statistics. To circumvent these issues,…

Machine Learning · Computer Science 2017-06-12 Santiago Pascual , Antonio Bonafonte , Joan Serrà

Recently, audio-visual separation approaches have taken advantage of the natural synchronization between the two modalities to boost audio source separation performance. They extracted high-level semantics from visual inputs as the guidance…

Sound · Computer Science 2024-07-08 Shentong Mo , Yapeng Tian

We propose WaveTrainerFit, a neural vocoder that performs high-quality waveform generation from data-driven features such as SSL features. WaveTrainerFit builds upon the WaveFit vocoder, which integrates diffusion model and generative…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-06 Hien Ohnaka , Yuma Shirahata , Masaya Kawamura

Enhancing speech quality under adverse SNR conditions remains a significant challenge for discriminative deep neural network (DNN)-based approaches. In this work, we propose DisCoGAN, which is a time-frequency-domain generative adversarial…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-18 Shrishti Saha Shetu , Emanuël A. P. Habets , Andreas Brendel

This paper focuses on using voice conversion (VC) to improve the speech intelligibility of surgical patients who have had parts of their articulators removed. Due to the difficulty of data collection, VC without parallel data is highly…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-26 Li-Wei Chen , Hung-Yi Lee , Yu Tsao

When trained on multimodal image datasets, normal Generative Adversarial Networks (GANs) are usually outperformed by class-conditional GANs and ensemble GANs, but conditional GANs is restricted to labeled datasets and ensemble GANs lack…

Computer Vision and Pattern Recognition · Computer Science 2019-01-29 Haifeng Shi , Guanyu Cai , Yuqin Wang , Shaohua Shang , Lianghua He

Recently, realistic data augmentation using neural networks especially generative neural networks (GAN) has achieved outstanding results. The communities main research focus is visual image processing. However, automotive cars and robots…

Computer Vision and Pattern Recognition · Computer Science 2019-02-27 Maximilian Pöpperl , Raghavendra Gulagundi , Senthil Yogamani , Stefan Milz

In this paper we demonstrate that it is possible to generate more meaningful electroencephalography (EEG) features from raw EEG features using generative adversarial networks (GAN) to improve the performance of EEG based continuous speech…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-03 Gautam Krishna , Co Tran , Mason Carnahan , Ahmed Tewfik

To effectively process impulse noise for narrowband powerline communications (NB-PLCs) transceivers, capturing comprehensive statistics of nonperiodic asynchronous impulsive noise (APIN) is a critical task. However, existing mathematical…

Signal Processing · Electrical Eng. & Systems 2025-10-30 Ying-Ren Chien , Po-Heng Chou , You-Jie Peng , Chun-Yuan Huang , Hen-Wai Tsao , Yu Tsao

Flow matching offers a robust and stable approach to training diffusion models. However, directly applying flow matching to neural vocoders can result in subpar audio quality. In this work, we present WaveFM, a reparameterized flow matching…

Sound · Computer Science 2025-03-24 Tianze Luo , Xingchen Miao , Wenbo Duan

Noise modeling lies in the heart of many image processing tasks. However, existing deep learning methods for noise modeling generally require clean and noisy image pairs for model training; these image pairs are difficult to obtain in many…

Computer Vision and Pattern Recognition · Computer Science 2020-06-05 Hanshu Yan , Xuan Chen , Vincent Y. F. Tan , Wenhan Yang , Joe Wu , Jiashi Feng

Due to the slow convergence and poor tracking ability, conventional LMS-based adaptive algorithms are less capable of handling dynamic noises. Selective fixed-filter active noise control (SFANC) can significantly reduce response time by…

Systems and Control · Electrical Eng. & Systems 2023-06-21 Zhengding Luo , Dongyuan Shi , Xiaoyi Shen , Junwei Ji , Woon-Seng Gan