English

Language-specific Acoustic Boundary Learning for Mandarin-English Code-switching Speech Recognition

Sound 2023-06-09 v1

Abstract

Code-switching speech recognition (CSSR) transcribes speech that switches between multiple languages or dialects within a single sentence. The main challenge in this task is that different languages often have similar pronunciations, making it difficult for models to distinguish between them. In this paper, we propose a method for solving the CSSR task from the perspective of language-specific acoustic boundary learning. We introduce language-specific weight estimators (LSWE) to model acoustic boundary learning in different languages separately. Additionally, a non-autoregressive (NAR) decoder and a language change detection (LCD) module are employed to assist in training. Evaluated on the SEAME corpus, our method achieves a state-of-the-art mixed error rate (MER) of 16.29% and 22.81% on the test_man and test_sge sets. We also demonstrate the effectiveness of our method on a 9000-hour in-house meeting code-switching dataset, where our method achieves a relatively 7.9% MER reduction.

Keywords

Cite

@article{arxiv.2306.05279,
  title  = {Language-specific Acoustic Boundary Learning for Mandarin-English Code-switching Speech Recognition},
  author = {Zhiyun Fan and Linhao Dong and Chen Shen and Zhenlin Liang and Jun Zhang and Lu Lu and Zejun Ma},
  journal= {arXiv preprint arXiv:2306.05279},
  year   = {2023}
}
R2 v1 2026-06-28T11:00:08.114Z