English

VoiceX: A Text-To-Speech Framework for Custom Voices

Human-Computer Interaction 2024-08-23 v1 Sound Audio and Speech Processing

Abstract

Modern TTS systems are capable of creating highly realistic and natural-sounding speech. Despite these developments, the process of customizing TTS voices remains a complex task, mostly requiring the expertise of specialists within the field. One reason for this is the utilization of deep learning models, which are characterized by their expansive, non-interpretable parameter spaces, restricting the feasibility of manual customization. In this paper, we present a novel human-in-the-loop paradigm based on an evolutionary algorithm for directly interacting with the parameter space of a neural TTS model. We integrated our approach into a user-friendly graphical user interface that allows users to efficiently create original voices. Those voices can then be used with the backbone TTS model, for which we provide a Python API. Further, we present the results of a user study exploring the capabilities of VoiceX. We show that VoiceX is an appropriate tool for creating individual, custom voices.

Keywords

Cite

@article{arxiv.2408.12170,
  title  = {VoiceX: A Text-To-Speech Framework for Custom Voices},
  author = {Silvan Mertes and Daksitha Withanage Don and Otto Grothe and Johanna Kuch and Ruben Schlagowski and Elisabeth André},
  journal= {arXiv preprint arXiv:2408.12170},
  year   = {2024}
}
R2 v1 2026-06-28T18:20:27.138Z