English

Generative Adversarial Training for Text-to-Speech Synthesis Based on Raw Phonetic Input and Explicit Prosody Modelling

Machine Learning 2023-10-17 v1

Abstract

We describe an end-to-end speech synthesis system that uses generative adversarial training. We train our Vocoder for raw phoneme-to-audio conversion, using explicit phonetic, pitch and duration modeling. We experiment with several pre-trained models for contextualized and decontextualized word embeddings and we introduce a new method for highly expressive character voice matching, based on discreet style tokens.

Keywords

Cite

@article{arxiv.2310.09636,
  title  = {Generative Adversarial Training for Text-to-Speech Synthesis Based on Raw Phonetic Input and Explicit Prosody Modelling},
  author = {Tiberiu Boros and Stefan Daniel Dumitrescu and Ionut Mironica and Radu Chivereanu},
  journal= {arXiv preprint arXiv:2310.09636},
  year   = {2023}
}
R2 v1 2026-06-28T12:50:44.411Z