English
Related papers

Related papers: High-quality nonparallel voice conversion based on…

200 papers

Besides the well-known classification task, these days neural networks are frequently being applied to generate or transform data, such as images and audio signals. In such tasks, the conventional loss functions like the mean squared error…

Generative Adversarial Networks (GANs) have become exceedingly popular in a wide range of data-driven research fields, due in part to their success in image generation. Their ability to generate new samples, often from only a small amount…

Computation and Language · Computer Science 2019-03-19 Thomas Wiest , Nicholas Cummins , Alice Baird , Simone Hantke , Judith Dineley , Björn Schuller

Current speaker recognition technology provides great performance with the x-vector approach. However, performance decreases when the evaluation domain is different from the training domain, an issue usually addressed with domain adaptation…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-29 Phani Sankar Nidadavolu , Saurabh Kataria , Jesús Villalba , Najim Dehak

This paper presents an adversarial learning method for recognition-synthesis based non-parallel voice conversion. A recognizer is used to transform acoustic features into linguistic representations while a synthesizer recovers output…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Jing-Xuan Zhang , Zhen-Hua Ling , Li-Rong Dai

Popular neural network-based speech enhancement systems operate on the magnitude spectrogram and ignore the phase mismatch between the noisy and clean speech signals. Conditional generative adversarial networks (cGANs) show promise in…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-21 Deepak Baby

Singing voice conversion aims to convert singer's voice from source to target without changing singing content. Parallel training data is typically required for the training of singing voice conversion system, that is however not practical…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-04 Junchen Lu , Kun Zhou , Berrak Sisman , Haizhou Li

In this paper, a new semi-supervised deep multiple-input multiple-output (MIMO) detection approach using a cycle-consistent generative adversarial network (CycleGAN) is proposed for communication systems without any prior knowledge of…

Signal Processing · Electrical Eng. & Systems 2023-04-24 Hongzhi Zhu , Yongliang Guo , Wei Xu , Xiaohu You

Numerous voice conversion (VC) techniques have been proposed for the conversion of voices among different speakers. Although good quality of the converted speech can be observed when VC is applied in a clean environment, the quality…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-20 Yun-Ju Chan , Chiang-Jen Peng , Syu-Siang Wang , Hsin-Min Wang , Yu Tsao , Tai-Shih Chi

Generative Adversarial Networks (GANs) have facilitated a new direction to tackle the image-to-image transformation problem. Different GANs use generator and discriminator networks with different losses in the objective function. Still…

Computer Vision and Pattern Recognition · Computer Science 2021-11-30 Kancharagunta Kishan Babu , Shiv Ram Dubey

We propose a new framework to improve automatic speech recognition (ASR) systems in resource-scarce environments using a generative adversarial network (GAN) operating on acoustic input features. The GAN is used to enhance the features of…

Sound · Computer Science 2022-10-07 Walter Heymans , Marelie H. Davel , Charl van Heerden

This paper evaluates the effectiveness of a Cycle-GAN based voice converter (VC) on four speaker identification (SID) systems and an automated speech recognition (ASR) system for various purposes. Audio samples converted by the VC model are…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-30 Gokce Keskin , Tyler Lee , Cory Stephenson , Oguz H. Elibol

In this paper, we present a description of the baseline system of Voice Conversion Challenge (VCC) 2020 with a cyclic variational autoencoder (CycleVAE) and Parallel WaveGAN (PWG), i.e., CycleVAEPWG. CycleVAE is a nonparallel VAE-based…

Sound · Computer Science 2020-10-12 Patrick Lumban Tobing , Yi-Chiao Wu , Tomoki Toda

In this paper, we propose a non-parallel any-to-many voice conversion (VC) method termed VoiceGrad. Inspired by WaveGrad, a recently introduced novel waveform generation method, VoiceGrad is based upon the concepts of score matching and…

Sound · Computer Science 2024-03-12 Hirokazu Kameoka , Takuhiro Kaneko , Kou Tanaka , Nobukatsu Hojo , Shogo Seki

In this paper, we present an open-source software for developing a nonparallel voice conversion (VC) system named crank. Although we have released an open-source VC software based on the Gaussian mixture model named sprocket in the last VC…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-05 Kazuhiro Kobayashi , Wen-Chin Huang , Yi-Chiao Wu , Patrick Lumban Tobing , Tomoki Hayashi , Tomoki Toda

Getting rid of the fundamental limitations in fitting to the paired training data, recent unsupervised low-light enhancement methods excel in adjusting illumination and contrast of images. However, for unsupervised low light enhancement,…

Computer Vision and Pattern Recognition · Computer Science 2022-07-05 Zhangkai Ni , Wenhan Yang , Hanli Wang , Shiqi Wang , Lin Ma , Sam Kwong

Non-parallel training is a difficult but essential task for DNN-based speech enhancement methods, for the lack of adequate noisy and paired clean speech corpus in many real scenarios. In this paper, we propose a novel adaptive…

Sound · Computer Science 2021-09-15 Guochen Yu , Yutian Wang , Chengshi Zheng , Hui Wang , Qin Zhang

The Generative Adversarial Networks (GANs) have demonstrated impressive performance for data synthesis, and are now used in a wide range of computer vision tasks. In spite of this success, they gained a reputation for being difficult to…

Machine Learning · Statistics 2017-12-07 Tatjana Chavdarova , François Fleuret

Emotional voice conversion aims to convert the spectrum and prosody to change the emotional patterns of speech, while preserving the speaker identity and linguistic content. Many studies require parallel speech data between different…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-27 Kun Zhou , Berrak Sisman , Haizhou Li

Speaking rate refers to the average number of phonemes within some unit time, while the rhythmic patterns refer to duration distributions for realizations of different phonemes within different phonetic structures. Both are key components…

Sound · Computer Science 2018-08-10 Cheng-chieh Yeh , Po-chun Hsu , Ju-chieh Chou , Hung-yi Lee , Lin-shan Lee

Nonparallel multi-domain voice conversion methods such as the StarGAN-VCs have been widely applied in many scenarios. However, the training of these models usually poses a challenge due to their complicated adversarial network…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-28 Shijing Si , Jianzong Wang , Xulong Zhang , Xiaoyang Qu , Ning Cheng , Jing Xiao