English

Multilingual Bottleneck Features for Improving ASR Performance of Code-Switched Speech in Under-Resourced Languages

Audio and Speech Processing 2020-11-09 v1 Computation and Language Machine Learning Sound

Abstract

In this work, we explore the benefits of using multilingual bottleneck features (mBNF) in acoustic modelling for the automatic speech recognition of code-switched (CS) speech in African languages. The unavailability of annotated corpora in the languages of interest has always been a primary challenge when developing speech recognition systems for this severely under-resourced type of speech. Hence, it is worthwhile to investigate the potential of using speech corpora available for other better-resourced languages to improve speech recognition performance. To achieve this, we train a mBNF extractor using nine Southern Bantu languages that form part of the freely available multilingual NCHLT corpus. We append these mBNFs to the existing MFCCs, pitch features and i-vectors to train acoustic models for automatic speech recognition (ASR) in the target code-switched languages. Our results show that the inclusion of the mBNF features leads to clear performance improvements over a baseline trained without the mBNFs for code-switched English-isiZulu, English-isiXhosa, English-Sesotho and English-Setswana speech.

Keywords

Cite

@article{arxiv.2011.03118,
  title  = {Multilingual Bottleneck Features for Improving ASR Performance of Code-Switched Speech in Under-Resourced Languages},
  author = {Trideba Padhi and Astik Biswas and Febe De Wet and Ewald van der Westhuizen and Thomas Niesler},
  journal= {arXiv preprint arXiv:2011.03118},
  year   = {2020}
}

Comments

In Proceedings of The First Workshop on Speech Technologies for Code-Switching in Multilingual Communities

R2 v1 2026-06-23T19:57:03.542Z