English

SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR

Sound 2024-06-27 v1 Machine Learning Audio and Speech Processing

Abstract

In this work, we propose a Switch-Conformer-based MoE system named SC-MoE for unified streaming and non-streaming code-switching (CS) automatic speech recognition (ASR), where we design a streaming MoE layer consisting of three language experts, which correspond to Mandarin, English, and blank, respectively, and equipped with a language identification (LID) network with a Connectionist Temporal Classification (CTC) loss as a router in the encoder of SC-MoE to achieve a real-time streaming CS ASR system. To further utilize the language information embedded in text, we also incorporate MoE layers into the decoder of SC-MoE. In addition, we introduce routers into every MoE layer of the encoder and the decoder and achieve better recognition performance. Experimental results show that the SC-MoE significantly improves CS ASR performances over baseline with comparable computational efficiency.

Keywords

Cite

@article{arxiv.2406.18021,
  title  = {SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR},
  author = {Shuaishuai Ye and Shunfei Chen and Xinhui Hu and Xinkang Xu},
  journal= {arXiv preprint arXiv:2406.18021},
  year   = {2024}
}

Comments

Accepted by InterSpeech 2024; 5 pages, 2 figures