English
Related papers

Related papers: Towards Improved Zero-shot Voice Conversion with C…

200 papers

Conflicting objectives present a considerable challenge in interleaving multi-task learning, necessitating the need for meticulous design and balance to ensure effective learning of a representative latent data space across all tasks…

Machine Learning · Computer Science 2025-01-17 Noelle Y. L. Wong , Eng Yeow Cheu , Zhonglin Chiam , Dipti Srinivasan

We propose a new unsupervised model for mapping a variable-duration speech segment to a fixed-dimensional representation. The resulting acoustic word embeddings can form the basis of search, discovery, and indexing systems for low- and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-07 Puyuan Peng , Herman Kamper , Karen Livescu

Tools to generate high quality synthetic speech signal that is perceptually indistinguishable from speech recorded from human speakers are easily available. Several approaches have been proposed for detecting synthetic speech. Many of these…

The human voice conveys not just words but also emotional states and individuality. Emotional voice conversion (EVC) modifies emotional expressions while preserving linguistic content and speaker identity, improving applications like…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-29 Hsing-Hang Chou , Yun-Shao Lin , Ching-Chin Sung , Yu Tsao , Chi-Chun Lee

Emotional voice conversion (VC) aims to convert a neutral voice to an emotional (e.g. happy) one while retaining the linguistic information and speaker identity. We note that the decoupling of emotional features from other speech…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-05 Zhaojie Luo , Shoufeng Lin , Rui Liu , Jun Baba , Yuichiro Yoshikawa , Ishiguro Hiroshi

Recently, a complex variational autoencoder (VAE)-based single-channel speech enhancement system based on the DCCRN architecture has been proposed. In this system, a noise suppression VAE (NSVAE) learns to extract clean speech…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-03 Jiatong Li , Simon Doclo

A variational autoencoder (VAE) is a probabilistic machine learning framework for posterior inference that projects an input set of high-dimensional data to a lower-dimensional, latent space. The latent space learned with a VAE offers…

Machine Learning · Computer Science 2022-11-16 Rafael Pastrana

Zero-shot voice conversion (VC) aims to transfer the timbre from the source speaker to an arbitrary unseen speaker while preserving the original linguistic content. Despite recent advancements in zero-shot VC using language model-based or…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-11 Jixun Yao , Yuguang Yang , Yu Pan , Ziqian Ning , Jiaohao Ye , Hongbin Zhou , Lei Xie

Denoising autoencoders (DAE) are trained to reconstruct their clean inputs with noise injected at the input level, while variational autoencoders (VAE) are trained with noise injected in their stochastic hidden layer, with a regularizer…

Machine Learning · Computer Science 2016-01-05 Daniel Jiwoong Im , Sungjin Ahn , Roland Memisevic , Yoshua Bengio

Learning disentangled representations of real-world data is a challenging open problem. Most previous methods have focused on either supervised approaches which use attribute labels or unsupervised approaches that manipulate the…

Computation and Language · Computer Science 2021-01-26 Vikash Balasubramanian , Ivan Kobyzev , Hareesh Bahuleyan , Ilya Shapiro , Olga Vechtomova

In this paper, we present a multimodal and dynamical VAE (MDVAE) applied to unsupervised audio-visual speech representation learning. The latent space is structured to dissociate the latent dynamical factors that are shared between the…

Sound · Computer Science 2024-02-21 Samir Sadok , Simon Leglaive , Laurent Girin , Xavier Alameda-Pineda , Renaud Séguier

Learning disentanglement aims at finding a low dimensional representation which consists of multiple explanatory and generative factors of the observational data. The framework of variational autoencoder (VAE) is commonly used to…

Machine Learning · Computer Science 2023-12-20 Mengyue Yang , Furui Liu , Zhitang Chen , Xinwei Shen , Jianye Hao , Jun Wang

Voice conversion (VC) has made progress in feature disentanglement, but it is still difficult to balance timbre and content information. This paper evaluates the pre-trained model features commonly used in voice conversion, and proposes an…

Sound · Computer Science 2025-04-09 Wenyu Wang , Yiquan Zhou , Jihua Zhu , Hongwu Ding , Jiacheng Xu , Shihao Li

We present a preliminary study on an end-to-end variational autoencoder (VAE) for sound morphing. Two VAE variants are compared: VAE with dilation layers (DC-VAE) and VAE only with regular convolutional layers (CC-VAE). We combine the…

Machine Learning · Computer Science 2020-11-20 Matteo Lionello , Hendrik Purwins

We propose SelfVC, a training strategy to iteratively improve a voice conversion model with self-synthesized examples. Previous efforts on voice conversion focus on factorizing speech into explicitly disentangled representations that…

This study addresses the problem of unsupervised subword unit discovery from untranscribed speech. It forms the basis of the ultimate goal of ZeroSpeech 2019, building text-to-speech systems without text labels. In this work, unit discovery…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-29 Siyuan Feng , Tan Lee , Zhiyuan Peng

As we enter the era of machine learning characterized by an overabundance of data, discovery, organization, and interpretation of the data in an unsupervised manner becomes a critical need. One promising approach to this endeavour is the…

Machine Learning · Computer Science 2022-10-24 Vaishnavi Patil , Matthew Evanusa , Joseph JaJa

Nowadays, as more and more systems achieve good performance in traditional voice conversion (VC) tasks, people's attention gradually turns to VC tasks under extreme conditions. In this paper, we propose a novel method for zero-shot voice…

Sound · Computer Science 2023-04-04 Haozhe Zhang , Zexin Cai , Xiaoyi Qin , Ming Li

Zero-shot voice conversion (VC) aims to convert the original speaker's timbre to any target speaker while keeping the linguistic content. Current mainstream zero-shot voice conversion approaches depend on pre-trained recognition models to…

Sound · Computer Science 2024-12-04 Yuke Li , Xinfa Zhu , Hanzhao Li , JiXun Yao , WenJie Tian , XiPeng Yang , YunLin Chen , Zhifei Li , Lei Xie

Recently, a generative variational autoencoder (VAE) has been proposed for speech enhancement to model speech statistics. However, this approach only uses clean speech in the training phase, making the estimation particularly sensitive to…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-18 Huajian Fang , Guillaume Carbajal , Stefan Wermter , Timo Gerkmann