BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
Abstract
We introduce a text-to-speech (TTS) model called BASE TTS, which stands for ig daptive treamable TTS with mergent abilities. BASE TTS is the largest TTS model to-date, trained on 100K hours of public domain speech data, achieving a new state-of-the-art in speech naturalness. It deploys a 1-billion-parameter autoregressive Transformer that converts raw texts into discrete codes ("speechcodes") followed by a convolution-based decoder which converts these speechcodes into waveforms in an incremental, streamable manner. Further, our speechcodes are built using a novel speech tokenization technique that features speaker ID disentanglement and compression with byte-pair encoding. Echoing the widely-reported "emergent abilities" of large language models when trained on increasing volume of data, we show that BASE TTS variants built with 10K+ hours and 500M+ parameters begin to demonstrate natural prosody on textually complex sentences. We design and share a specialized dataset to measure these emergent abilities for text-to-speech. We showcase state-of-the-art naturalness of BASE TTS by evaluating against baselines that include publicly available large-scale text-to-speech systems: YourTTS, Bark and TortoiseTTS. Audio samples generated by the model can be heard at https://amazon-ltts-paper.com/.
Cite
@article{arxiv.2402.08093,
title = {BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data},
author = {Mateusz Łajszczak and Guillermo Cámbara and Yang Li and Fatih Beyhan and Arent van Korlaar and Fan Yang and Arnaud Joly and Álvaro Martín-Cortinas and Ammar Abbas and Adam Michalski and Alexis Moinet and Sri Karlapati and Ewa Muszyńska and Haohan Guo and Bartosz Putrycz and Soledad López Gambino and Kayeon Yoo and Elena Sokolova and Thomas Drugman},
journal= {arXiv preprint arXiv:2402.08093},
year = {2024}
}
Comments
v1.1 (fixed typos)