Controllable Generation of Artificial Speaker Embeddings through Discovery of Principal Directions
Abstract
Customizing voice and speaking style in a speech synthesis system with intuitive and fine-grained controls is challenging, given that little data with appropriate labels is available. Furthermore, editing an existing human's voice also comes with ethical concerns. In this paper, we propose a method to generate artificial speaker embeddings that cannot be linked to a real human while offering intuitive and fine-grained control over the voice and speaking style of the embeddings, without requiring any labels for speaker or style. The artificial and controllable embeddings can be fed to a speech synthesis system, conditioned on embeddings of real humans during training, without sacrificing privacy during inference.
Keywords
Cite
@article{arxiv.2310.17502,
title = {Controllable Generation of Artificial Speaker Embeddings through Discovery of Principal Directions},
author = {Florian Lux and Pascal Tilli and Sarina Meyer and Ngoc Thang Vu},
journal= {arXiv preprint arXiv:2310.17502},
year = {2023}
}
Comments
Published at ISCA Interspeech 2023 https://www.isca-speech.org/archive/interspeech_2023/lux23_interspeech.html