English

WAXAL: A Large-Scale Multilingual African Language Speech Corpus

Audio and Speech Processing 2026-03-03 v3 Artificial Intelligence Computation and Language

Abstract

The advancement of speech technology has predominantly favored high-resource languages, creating a significant digital divide for speakers of most Sub-Saharan African languages. To address this gap, we introduce WAXAL, a large-scale, openly accessible speech dataset for 24 languages representing over 100 million speakers. The collection consists of two main components: an Automated Speech Recognition (ASR) dataset containing approximately 1,250 hours of transcribed, natural speech from a diverse range of speakers, and a Text-to-Speech (TTS) dataset with around 235 hours of high-quality, single-speaker recordings reading phonetically balanced scripts. This paper details our methodology for data collection, annotation, and quality control, which involved partnerships with four African academic and community organizations. We provide a detailed statistical overview of the dataset and discuss its potential limitations and ethical considerations. The WAXAL datasets are released at https://huggingface.co/datasets/google/WaxalNLP under the permissive CC-BY-4.0 license to catalyze research, enable the development of inclusive technologies, and serve as a vital resource for the digital preservation of these languages.

Cite

@article{arxiv.2602.02734,
  title  = {WAXAL: A Large-Scale Multilingual African Language Speech Corpus},
  author = {Abdoulaye Diack and Perry Nelson and Kwaku Agbesi and Angela Nakalembe and MohamedElfatih MohamedKhair and Vusumuzi Dube and Tavonga Siyavora and Subhashini Venugopalan and Jason Hickey and Uche Okonkwo and Abhishek Bapna and Isaac Wiafe and Raynard Dodzi Helegah and Elikem Doe Atsakpo and Charles Nutrokpor and Fiifi Baffoe Payin Winful and Kafui Kwashie Solaga and Jamal-Deen Abdulai and Akon Obu Ekpezu and Audace Niyonkuru and Samuel Rutunda and Boris Ishimwe and Michael Melese and Engineer Bainomugisha and Joyce Nakatumba-Nabende and Andrew Katumba and Claire Babirye and Jonathan Mukiibi and Vincent Kimani and Samuel Kibacia and James Maina and Fridah Emmah and Ahmed Ibrahim Shekarau and Ibrahim Shehu Adamu and Yusuf Abdullahi and Howard Lakougna and Bob MacDonald and Hadar Shemtov and Aisha Walcott-Bryant and Moustapha Cisse and Avinatan Hassidim and Jeff Dean and Yossi Matias},
  journal= {arXiv preprint arXiv:2602.02734},
  year   = {2026}
}

Comments

Initial dataset release with added TTS, some more to come

R2 v1 2026-07-01T09:32:55.259Z