English

Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition

Computation and Language 2026-06-30 v1

Abstract

Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits practical use in education and public services. We addressed this gap with a tone conditioned curriculum framework for 6 Southern Bantu languages that combined hybrid difficulty scoring, gated adapters driven by tonal statistics and staged curriculum training. We trained on a community corpus and tested transfer to NCHLT to measure robustness beyond matched evaluation. Results revealed clear interactions between architecture and language, with W2V-BERT outperforming Whisper on Nguni languages by 3 to 4 WER points whilst Whisper performed better on Sotho-Tswana languages. W2V-BERT with tone conditioning reached 28.41% average WER across datasets and 23.79% on Xitsonga transfer. No single model suited all 6 languages, so deployment should pair model selection per language with validation across corpora.

Cite

@article{arxiv.2606.31642,
  title  = {Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition},
  author = {Kesego Mokgosi and Vukosi Marivate and Sitwala Mundia and Unarine Netshifhefhe and Tsholofelo Hope Mogale and Thapelo Sindane},
  journal= {arXiv preprint arXiv:2606.31642},
  year   = {2026}
}