One Model to Pronounce Them All: Multilingual Grapheme-to-Phoneme Conversion With a Transformer Ensemble
Computation and Language
2020-06-25 v1
Abstract
The task of grapheme-to-phoneme (G2P) conversion is important for both speech recognition and synthesis. Similar to other speech and language processing tasks, in a scenario where only small-sized training data are available, learning G2P models is challenging. We describe a simple approach of exploiting model ensembles, based on multilingual Transformers and self-training, to develop a highly effective G2P solution for 15 languages. Our models are developed as part of our participation in the SIGMORPHON 2020 Shared Task 1 focused at G2P. Our best models achieve 14.99 word error rate (WER) and 3.30 phoneme error rate (PER), a sizeable improvement over the shared task competitive baselines.
Keywords
Cite
@article{arxiv.2006.13343,
title = {One Model to Pronounce Them All: Multilingual Grapheme-to-Phoneme Conversion With a Transformer Ensemble},
author = {Kaili Vesik and Muhammad Abdul-Mageed and Miikka Silfverberg},
journal= {arXiv preprint arXiv:2006.13343},
year = {2020}
}
Comments
7 pages, submitted to SIGMORPHON 2020 Shared Task 1