English
Related papers

Related papers: Voice Conversion with Diverse Intonation using Con…

200 papers

We propose a cross-lingual neural codec language model, VALL-E X, for cross-lingual speech synthesis. Specifically, we extend VALL-E and train a multi-lingual conditional codec language model to predict the acoustic token sequences of the…

Computation and Language · Computer Science 2023-03-08 Ziqiang Zhang , Long Zhou , Chengyi Wang , Sanyuan Chen , Yu Wu , Shujie Liu , Zhuo Chen , Yanqing Liu , Huaming Wang , Jinyu Li , Lei He , Sheng Zhao , Furu Wei

Non-parallel voice conversion aims to convert voice from a source domain to a target domain without paired training data. Cycle-Consistent Generative Adversarial Networks (CycleGAN) and Variational Autoencoders (VAE) have been used for this…

Sound · Computer Science 2025-10-16 Maharnab Saikia

Many factors influence speech yielding different renditions of a given sentence. Generative models, such as variational autoencoders (VAEs), capture this variability and allow multiple renditions of the same sentence via sampling. The…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-21 Penny Karanasou , Sri Karlapati , Alexis Moinet , Arnaud Joly , Ammar Abbas , Simon Slangen , Jaime Lorenzo Trueba , Thomas Drugman

Accent conversion (AC) transforms a non-native speaker's accent into a native accent while maintaining the speaker's voice timbre. In this paper, we propose approaches to improving accent conversion applicability, as well as quality. First…

Computation and Language · Computer Science 2020-05-20 Wenjie Li , Benlai Tang , Xiang Yin , Yushi Zhao , Wei Li , Kang Wang , Hao Huang , Yuxuan Wang , Zejun Ma

In English, prosody adds a broad range of information to segment sequences, from information structure (e.g. contrast) to stylistic variation (e.g. expression of emotion). However, when learning to control prosody in text-to-speech voices,…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-04 Zack Hodari , Catherine Lai , Simon King

This paper proposes a voice conversion (VC) method using sequence-to-sequence (seq2seq or S2S) learning, which flexibly converts not only the voice characteristics but also the pitch contour and duration of input speech. The proposed…

Sound · Computer Science 2020-10-08 Hirokazu Kameoka , Kou Tanaka , Damian Kwasny , Takuhiro Kaneko , Nobukatsu Hojo

In this paper, we present a deep generative model based method to generate diverse human motion interpolation results. We resort to the Conditional Variational Auto-Encoder (CVAE) to learn human motion conditioned on a pair of given start…

Computer Vision and Pattern Recognition · Computer Science 2021-11-15 Chunzhi Gu , Shuofeng Zhao , Chao Zhang

State-of-the-art Variational Auto-Encoders (VAEs) for learning disentangled latent representations give impressive results in discovering features like pitch, pause duration, and accent in speech data, leading to highly controllable…

Sound · Computer Science 2021-05-11 Shakti Kumar , Jithin Pradeep , Hussain Zaidi

Singing voice conversion is to convert the source singing voice into the target singing voice except for the content. Currently, flow-based models can complete the task of voice conversion, but they struggle to effectively extract latent…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-10 Hui Li , Hongyu Wang , Zhijin Chen , Bohan Sun , Bo Li

Generative models have been successfully applied to image style transfer and domain translation. However, there is still a wide gap in the quality of results when learning such tasks on musical audio. Furthermore, most translation models…

Sound · Computer Science 2018-10-02 Adrien Bitton , Philippe Esling , Axel Chemla-Romeu-Santos

In this paper, we propose a new approach to pathological speech synthesis. Instead of using healthy speech as a source, we customise an existing pathological speech sample to a new speaker's voice characteristics. This approach alleviates…

Selective manipulation of data attributes using deep generative models is an active area of research. In this paper, we present a novel method to structure the latent space of a Variational Auto-Encoder (VAE) to encode different…

Machine Learning · Computer Science 2020-07-30 Ashis Pati , Alexander Lerch

In this paper we present a new implementation of a Variational Autoencoder (VAE) for the calibration of sensors. We propose that the VAE can be used to calibrate sensor data by training the latent space as a calibration output. We discuss…

Machine Learning · Computer Science 2025-11-04 Travis Barrett , Amit Kumar Mishra , Joyce Mwangama

Substantial improvements have been achieved in recent years in voice conversion, which converts the speaker characteristics of an utterance into those of another speaker without changing the linguistic content of the utterance. Nonetheless,…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-05 Chien-yu Huang , Yist Y. Lin , Hung-yi Lee , Lin-shan Lee

Posterior collapse plagues VAEs for text, especially for conditional text generation with strong autoregressive decoders. In this work, we address this problem in variational neural machine translation by explicitly promoting mutual…

Computation and Language · Computer Science 2019-09-23 Arya D. McCarthy , Xian Li , Jiatao Gu , Ning Dong

The goal of voice conversion is to transform source speech into a target voice, keeping the content unchanged. In this paper, we focus on self-supervised representation learning for voice conversion. Specifically, we compare discrete and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-09 Benjamin van Niekerk , Marc-André Carbonneau , Julian Zaïdi , Mathew Baas , Hugo Seuté , Herman Kamper

Recent advancements in Text-to-Speech (TTS) systems have enabled the generation of natural and expressive speech from textual input. Accented TTS aims to enhance user experience by making the synthesized speech more relatable to minority…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-18 Jan Melechovsky , Ambuj Mehrish , Berrak Sisman , Dorien Herremans

Variational autoencoders (VAEs) are one class of generative probabilistic latent-variable models designed for inference based on known data. We develop three variations on VAEs by introducing a second parameterized encoder/decoder pair and,…

Machine Learning · Computer Science 2023-04-06 R. I. Cukier

Expressive voice conversion aims to transfer both speaker identity and expressive attributes from a target speech to a given source speech. In this work, we improve over a self-supervised, non-autoregressive framework with a conditional…

Sound · Computer Science 2025-06-05 Seymanur Akti , Tuan Nam Nguyen , Alexander Waibel

Recent advances in neural autoregressive models have improve the performance of speech synthesis (SS). However, as they lack the ability to model global characteristics of speech (such as speaker individualities or speaking styles),…

Computation and Language · Computer Science 2019-02-12 Kei Akuzawa , Yusuke Iwasawa , Yutaka Matsuo