English

Africa-Centric Self-Supervised Pre-Training for Multilingual Speech Representation in a Sub-Saharan Context

Computation and Language 2024-04-23 v3 Machine Learning Sound Audio and Speech Processing

Abstract

We present the first self-supervised multilingual speech model trained exclusively on African speech. The model learned from nearly 60 000 hours of unlabeled speech segments in 21 languages and dialects spoken in sub-Saharan Africa. On the SSA subset of the FLEURS-102 dataset, our approach based on a HuBERTbase_{base} (0.09B) architecture shows competitive results, for ASR downstream task, compared to the w2v-bert-51 (0.6B) pre-trained model proposed in the FLEURS benchmark, while being more efficient by using 7x less data and 6x less parameters. Furthermore, in the context of a LID downstream task, our approach outperforms FLEURS baselines accuracy by over 22\%.

Keywords

Cite

@article{arxiv.2404.02000,
  title  = {Africa-Centric Self-Supervised Pre-Training for Multilingual Speech Representation in a Sub-Saharan Context},
  author = {Antoine Caubrière and Elodie Gauthier},
  journal= {arXiv preprint arXiv:2404.02000},
  year   = {2024}
}

Comments

To appear in AfricaNLP 2024

R2 v1 2026-06-28T15:41:46.711Z