中文
相关论文

相关论文: Basis-MelGAN: Efficient Neural Vocoder Based on Au…

200 篇论文

We investigate the effectiveness of generative adversarial networks (GANs) for speech enhancement, in the context of improving noise robustness of automatic speech recognition (ASR) systems. Prior work demonstrates that GANs can effectively…

声音 · 计算机科学 2018-11-01 Chris Donahue , Bo Li , Rohit Prabhavalkar

Recent studies have shown that text-to-speech synthesis quality can be improved by using glottal vocoding. This refers to vocoders that parameterize speech into two parts, the glottal excitation and vocal tract, that occur in the human…

音频与语音处理 · 电气工程与系统科学 2019-03-15 Bajibabu Bollepalli , Lauri Juvela , Paavo Alku

This paper presents a deep learning-based approach for the spatio-temporal reconstruction of sound fields using Generative Adversarial Networks (GANs). The method utilises a plane wave basis and learns the underlying statistical…

音频与语音处理 · 电气工程与系统科学 2023-08-02 Xenofon Karakonstantis , Efren Fernandez-Grande

The quality of speech coded by transform coding is affected by various artefacts especially when bitrates to quantize the frequency components become too low. In order to mitigate these coding artefacts and enhance the quality of coded…

音频与语音处理 · 电气工程与系统科学 2022-02-01 Srikanth Korse , Nicola Pia , Kishan Gupta , Guillaume Fuchs

In this paper, we proposed AI-based audio coding using MFCC features in an adversarial setting. We combined a conventional encoder with an adversarial learning decoder to better reconstruct the original waveform. Since GAN gives implicit…

音频与语音处理 · 电气工程与系统科学 2023-10-24 Mohammad Reza Hasanabadi

The performance of speech processing models trained on clean speech drops significantly in noisy conditions. Training with noisy datasets alleviates the problem, but procuring such datasets is not always feasible. Noisy speech simulation…

声音 · 计算机科学 2023-05-23 Leander Melroy Maben , Zixun Guo , Chen Chen , Utkarsh Chudiwal , Chng Eng Siong

In this work, we further develop the conformer-based metric generative adversarial network (CMGAN) model for speech enhancement (SE) in the time-frequency (TF) domain. This paper builds on our previous work but takes a more in-depth look by…

声音 · 计算机科学 2024-05-07 Sherif Abdulatif , Ruizhe Cao , Bin Yang

Point clouds acquired from range scans are often sparse, noisy, and non-uniform. This paper presents a new point cloud upsampling network called PU-GAN, which is formulated based on a generative adversarial network (GAN), to learn a rich…

计算机视觉与模式识别 · 计算机科学 2019-07-26 Ruihui Li , Xianzhi Li , Chi-Wing Fu , Daniel Cohen-Or , Pheng-Ann Heng

Voice conversion is a method that allows for the transformation of speaking style while maintaining the integrity of linguistic information. There are many researchers using deep generative models for voice conversion tasks. Generative…

声音 · 计算机科学 2023-08-29 Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao

Synthesizing high quality saliency maps from noisy images is a challenging problem in computer vision and has many practical applications. Samples generated by existing techniques for saliency detection cannot handle the noise perturbations…

计算机视觉与模式识别 · 计算机科学 2019-04-03 Prerana Mukherjee , Manoj Sharma , Megh Makwana , Ajay Pratap Singh , Avinash Upadhyay , Akkshita Trivedi , Brejesh Lall , Santanu Chaudhury

Generative Adversarial Networks (GANs) have seen steep ascension to the peak of ML research zeitgeist in recent years. Mostly catalyzed by its success in the domain of image generation, the technique has seen wide range of adoption in a…

机器学习 · 统计学 2018-05-09 Aparna Balagopalan , Satya Gorti , Mathieu Ravaut , Raeid Saqur

Neural vocoders model the raw audio waveform and synthesize high-quality audio, but even the highly efficient ones, like MB-MelGAN and LPCNet, fail to run real-time on a low-end device like a smartglass. A pure digital signal processing…

声音 · 计算机科学 2024-01-22 Prabhav Agrawal , Thilo Koehler , Zhiping Xiu , Prashant Serai , Qing He

To effectively process impulse noise for narrowband powerline communications (NB-PLCs) transceivers, capturing comprehensive statistics of nonperiodic asynchronous impulsive noise (APIN) is a critical task. However, existing mathematical…

信号处理 · 电气工程与系统科学 2025-10-30 Ying-Ren Chien , Po-Heng Chou , You-Jie Peng , Chun-Yuan Huang , Hen-Wai Tsao , Yu Tsao

Enhancing speech quality under adverse SNR conditions remains a significant challenge for discriminative deep neural network (DNN)-based approaches. In this work, we propose DisCoGAN, which is a time-frequency-domain generative adversarial…

音频与语音处理 · 电气工程与系统科学 2024-10-18 Shrishti Saha Shetu , Emanuël A. P. Habets , Andreas Brendel

Separating two sources from an audio mixture is an important task with many applications. It is a challenging problem since only one signal channel is available for analysis. In this paper, we propose a novel framework for singing voice…

声音 · 计算机科学 2017-11-15 Zhe-Cheng Fan , Yen-Lin Lai , Jyh-Shing Roger Jang

While neural vocoders have made significant progress in high-fidelity speech synthesis, their application on polyphonic music has remained underexplored. In this work, we propose DisCoder, a neural vocoder that leverages a generative…

We present the first neural video compression method based on generative adversarial networks (GANs). Our approach significantly outperforms previous neural and non-neural video compression methods in a user study, setting a new…

图像与视频处理 · 电气工程与系统科学 2022-07-13 Fabian Mentzer , Eirikur Agustsson , Johannes Ballé , David Minnen , Nick Johnston , George Toderici

Speech enhancement at extremely low signal-to-noise ratio (SNR) condition is a very challenging problem and rarely investigated in previous works. This paper proposes a robust speech enhancement approach (UNetGAN) based on U-Net and…

音频与语音处理 · 电气工程与系统科学 2020-10-30 Xiang Hao , Xiangdong Su , Zhiyu Wang , Hui Zhang , Batushiren

Current two-stage TTS framework typically integrates an acoustic model with a vocoder -- the acoustic model predicts a low resolution intermediate representation such as Mel-spectrum while the vocoder generates waveform from the…

音频与语音处理 · 电气工程与系统科学 2021-06-23 Jian Cong , Shan Yang , Lei Xie , Dan Su

Popular neural network-based speech enhancement systems operate on the magnitude spectrogram and ignore the phase mismatch between the noisy and clean speech signals. Conditional generative adversarial networks (cGANs) show promise in…

音频与语音处理 · 电气工程与系统科学 2020-02-21 Deepak Baby