English

Language-Universal Phonetic Representation in Multilingual Speech Pretraining for Low-Resource Speech Recognition

Audio and Speech Processing 2023-05-22 v1 Computation and Language Sound

Abstract

We improve low-resource ASR by integrating the ideas of multilingual training and self-supervised learning. Concretely, we leverage an International Phonetic Alphabet (IPA) multilingual model to create frame-level pseudo labels for unlabeled speech, and use these pseudo labels to guide hidden-unit BERT (HuBERT) based speech pretraining in a phonetically-informed manner. The experiments on the Multilingual Speech (MLS) Corpus show that the proposed approach consistently outperforms the standard HuBERT on all the target languages. Moreover, on 3 of the 4 languages, comparing to the standard HuBERT, the approach performs better, meanwhile is able to save supervised training data by 1.5k hours (75%) at most. Our approach outperforms most of the state of the arts, with much less pretraining data in terms of hours and language diversity. Compared to XLSR-53 and a retraining based multilingual method, our approach performs better with full and limited finetuning data scenarios.

Keywords

Cite

@article{arxiv.2305.11569,
  title  = {Language-Universal Phonetic Representation in Multilingual Speech Pretraining for Low-Resource Speech Recognition},
  author = {Siyuan Feng and Ming Tu and Rui Xia and Chuanzeng Huang and Yuxuan Wang},
  journal= {arXiv preprint arXiv:2305.11569},
  year   = {2023}
}

Comments

Accepted for publication in INTERSPEECH 2023

R2 v1 2026-06-28T10:39:05.646Z