English

Google Crowdsourced Speech Corpora and Related Open-Source Resources for Low-Resource Languages and Dialects: An Overview

Computation and Language 2020-10-15 v1

Abstract

This paper presents an overview of a program designed to address the growing need for developing freely available speech resources for under-represented languages. At present we have released 38 datasets for building text-to-speech and automatic speech recognition applications for languages and dialects of South and Southeast Asia, Africa, Europe and South America. The paper describes the methodology used for developing such corpora and presents some of our findings that could benefit under-represented language communities.

Keywords

Cite

@article{arxiv.2010.06778,
  title  = {Google Crowdsourced Speech Corpora and Related Open-Source Resources for Low-Resource Languages and Dialects: An Overview},
  author = {Alena Butryna and Shan-Hui Cathy Chu and Isin Demirsahin and Alexander Gutkin and Linne Ha and Fei He and Martin Jansche and Cibu Johny and Anna Katanova and Oddur Kjartansson and Chenfang Li and Tatiana Merkulova and Yin May Oo and Knot Pipatsrisawat and Clara Rivera and Supheakmungkol Sarin and Pasindu de Silva and Keshan Sodimana and Richard Sproat and Theeraphol Wattanavekin and Jaka Aris Eko Wibawa},
  journal= {arXiv preprint arXiv:2010.06778},
  year   = {2020}
}

Comments

Appeared in 2019 UNESCO International Conference Language Technologies for All (LT4All): Enabling Linguistic Diversity and Multilingualism Worldwide, 4-6 December, Paris, France

R2 v1 2026-06-23T19:19:44.354Z