English

DistilXLSR: A Light Weight Cross-Lingual Speech Representation Model

Computation and Language 2023-06-05 v1 Sound Audio and Speech Processing

Abstract

Multilingual self-supervised speech representation models have greatly enhanced the speech recognition performance for low-resource languages, and the compression of these huge models has also become a crucial prerequisite for their industrial application. In this paper, we propose DistilXLSR, a distilled cross-lingual speech representation model. By randomly shuffling the phonemes of existing speech, we reduce the linguistic information and distill cross-lingual models using only English data. We also design a layer-jumping initialization method to fully leverage the teacher's pre-trained weights. Experiments on 2 kinds of teacher models and 15 low-resource languages show that our method can reduce the parameters by 50% while maintaining cross-lingual representation ability. Our method is proven to be generalizable to various languages/teacher models and has the potential to improve the cross-lingual performance of the English pre-trained models.

Keywords

Cite

@article{arxiv.2306.01303,
  title  = {DistilXLSR: A Light Weight Cross-Lingual Speech Representation Model},
  author = {Haoyu Wang and Siyuan Wang and Wei-Qiang Zhang and Jinfeng Bai},
  journal= {arXiv preprint arXiv:2306.01303},
  year   = {2023}
}

Comments

Accepted by INTERSPEECH 2023

R2 v1 2026-06-28T10:54:15.134Z