English

RyanSpeech: A Corpus for Conversational Text-to-Speech Synthesis

Computation and Language 2021-06-17 v1 Sound Audio and Speech Processing

Abstract

This paper introduces RyanSpeech, a new speech corpus for research on automated text-to-speech (TTS) systems. Publicly available TTS corpora are often noisy, recorded with multiple speakers, or lack quality male speech data. In order to meet the need for a high quality, publicly available male speech corpus within the field of speech recognition, we have designed and created RyanSpeech which contains textual materials from real-world conversational settings. These materials contain over 10 hours of a professional male voice actor's speech recorded at 44.1 kHz. This corpus's design and pipeline make RyanSpeech ideal for developing TTS systems in real-world applications. To provide a baseline for future research, protocols, and benchmarks, we trained 4 state-of-the-art speech models and a vocoder on RyanSpeech. The results show 3.36 in mean opinion scores (MOS) in our best model. We have made both the corpus and trained models for public use.

Keywords

Cite

@article{arxiv.2106.08468,
  title  = {RyanSpeech: A Corpus for Conversational Text-to-Speech Synthesis},
  author = {Rohola Zandie and Mohammad H. Mahoor and Julia Madsen and Eshrat S. Emamian},
  journal= {arXiv preprint arXiv:2106.08468},
  year   = {2021}
}
R2 v1 2026-06-24T03:14:41.363Z