中文
相关论文

相关论文: MaskCycleGAN-VC: Learning Non-parallel Voice Conve…

200 篇论文

In a typical voice conversion system, vocoder is commonly used for speech-to-features analysis and features-to-speech synthesis. However, vocoder can be a source of speech quality degradation. This paper presents a vocoder-free voice…

音频与语音处理 · 电气工程与系统科学 2019-09-18 Xiaohai Tian , Eng Siong Chng , Haizhou Li

In this paper, we present the voice conversion (VC) systems developed at Nagoya University (NU) for the Voice Conversion Challenge 2020 (VCC2020). We aim to determine the effectiveness of two recent significant technologies in VC:…

音频与语音处理 · 电气工程与系统科学 2020-10-12 Wen-Chin Huang , Patrick Lumban Tobing , Yi-Chiao Wu , Kazuhiro Kobayashi , Tomoki Toda

This paper proposes a new voice conversion (VC) task from human speech to dog-like speech while preserving linguistic information as an example of human to non-human creature voice conversion (H2NH-VC) tasks. Although most VC studies deal…

声音 · 计算机科学 2023-01-18 Kohei Suzuki , Shoki Sakamoto , Tadahiro Taniguchi , Hirokazu Kameoka

This paper proposes a voice conversion (VC) method based on a sequence-to-sequence (S2S) learning framework, which enables simultaneous conversion of the voice characteristics, pitch contour, and duration of input speech. We previously…

音频与语音处理 · 电气工程与系统科学 2020-11-10 Hirokazu Kameoka , Wen-Chin Huang , Kou Tanaka , Takuhiro Kaneko , Nobukatsu Hojo , Tomoki Toda

We introduce a novel sequence-to-sequence (seq2seq) voice conversion (VC) model based on the Transformer architecture with text-to-speech (TTS) pretraining. Seq2seq VC models are attractive owing to their ability to convert prosody. While…

音频与语音处理 · 电气工程与系统科学 2019-12-17 Wen-Chin Huang , Tomoki Hayashi , Yi-Chiao Wu , Hirokazu Kameoka , Tomoki Toda

Precise control over speech characteristics, such as pitch, duration, and speech rate, remains a significant challenge in the field of voice conversion. The ability to manipulate parameters like pitch and syllable rate is an important…

声音 · 计算机科学 2025-07-08 Mathilde Abrassart , Nicolas Obin , Axel Roebel

In recent years, neural vocoders have surpassed classical speech generation approaches in naturalness and perceptual quality of the synthesized speech. Computationally heavy models like WaveNet and WaveGlow achieve best results, while…

音频与语音处理 · 电气工程与系统科学 2021-02-15 Ahmed Mustafa , Nicola Pia , Guillaume Fuchs

For the lack of adequate paired noisy-clean speech corpus in many real scenarios, non-parallel training is a promising task for DNN-based speech enhancement methods. However, because of the severe mismatch between input and target speeches,…

声音 · 计算机科学 2022-02-15 Guochen Yu , Andong Li , Yutian Wang , Yinuo Guo , Hui Wang , Chengshi Zheng

An effective approach to non-parallel voice conversion (VC) is to utilize deep neural networks (DNNs), specifically variational auto encoders (VAEs), to model the latent structure of speech in an unsupervised manner. A previous study has…

音频与语音处理 · 电气工程与系统科学 2020-04-09 Wen-Chin Huang , Hsin-Te Hwang , Yu-Huai Peng , Yu Tsao , Hsin-Min Wang

This paper presents a low-latency real-time (LLRT) non-parallel voice conversion (VC) framework based on cyclic variational autoencoder (CycleVAE) and multiband WaveRNN with data-driven linear prediction (MWDLP). CycleVAE is a robust…

声音 · 计算机科学 2021-07-06 Patrick Lumban Tobing , Tomoki Toda

Domain adaptation plays an important role for speech recognition models, in particular, for domains that have low resources. We propose a novel generative model based on cyclic-consistent generative adversarial network (CycleGAN) for…

计算与语言 · 计算机科学 2018-07-11 Ehsan Hosseini-Asl , Yingbo Zhou , Caiming Xiong , Richard Socher

Voice conversion (VC) consists of digitally altering the voice of an individual to manipulate part of its content, primarily its identity, while maintaining the rest unchanged. Research in neural VC has accomplished considerable…

声音 · 计算机科学 2021-07-28 Laurent Benaroya , Nicolas Obin , Axel Roebel

With the increase in the availability of speech from varied domains, it is imperative to use such out-of-domain data to improve existing speech systems. Domain adaptation is a prominent pre-processing approach for this. We investigate it…

音频与语音处理 · 电气工程与系统科学 2021-04-06 Saurabh Kataria , Jesús Villalba , Piotr Żelasko , Laureano Moro-Velázquez , Najim Dehak

Mel-frequency filter bank (MFB) based approaches have the advantage of learning speech compared to raw spectrum since MFB has less feature size. However, speech generator with MFB approaches require additional vocoder that needs a huge…

音频与语音处理 · 电气工程与系统科学 2020-12-01 June-Woo Kim , Ho-Young Jung , Minho Lee

This paper describes a method based on a sequence-to-sequence learning (Seq2Seq) with attention and context preservation mechanism for voice conversion (VC) tasks. Seq2Seq has been outstanding at numerous tasks involving sequence modeling…

音频与语音处理 · 电气工程与系统科学 2018-11-13 Kou Tanaka , Hirokazu Kameoka , Takuhiro Kaneko , Nobukatsu Hojo

Singing voice conversion (SVC) aims to convert the voice of one singer to that of other singers while keeping the singing content and melody. On top of recent voice conversion works, we propose a novel model to steadily convert songs while…

声音 · 计算机科学 2020-10-29 Zhonghao Li , Benlai Tang , Xiang Yin , Yuan Wan , Ling Xu , Chen Shen , Zejun Ma

Foreign accent conversion (FAC) is a special application of voice conversion (VC) which aims to convert the accented speech of a non-native speaker to a native-sounding speech with the same speaker identity. FAC is difficult since the…

声音 · 计算机科学 2023-09-06 Wen-Chin Huang , Tomoki Toda

Numerous voice conversion (VC) techniques have been proposed for the conversion of voices among different speakers. Although good quality of the converted speech can be observed when VC is applied in a clean environment, the quality…

音频与语音处理 · 电气工程与系统科学 2023-01-20 Yun-Ju Chan , Chiang-Jen Peng , Syu-Siang Wang , Hsin-Min Wang , Yu Tsao , Tai-Shih Chi

In this paper, we integrate a simple non-parallel voice conversion (VC) system with a WaveNet (WN) vocoder and a proposed collapsed speech suppression technique. The effectiveness of WN as a vocoder for generating high-fidelity speech…

音频与语音处理 · 电气工程与系统科学 2020-04-08 Yi-Chiao Wu , Patrick Lumban Tobing , Kazuhiro Kobayashi , Tomoki Hayashi , Tomoki Toda

The voice conversion challenge is a bi-annual scientific event held to compare and understand different voice conversion (VC) systems built on a common dataset. In 2020, we organized the third edition of the challenge and constructed and…

音频与语音处理 · 电气工程与系统科学 2020-08-31 Yi Zhao , Wen-Chin Huang , Xiaohai Tian , Junichi Yamagishi , Rohan Kumar Das , Tomi Kinnunen , Zhenhua Ling , Tomoki Toda