English

Bilingual End-to-End ASR with Byte-Level Subwords

Computation and Language 2022-05-03 v1 Sound Audio and Speech Processing

Abstract

In this paper, we investigate how the output representation of an end-to-end neural network affects multilingual automatic speech recognition (ASR). We study different representations including character-level, byte-level, byte pair encoding (BPE), and byte-level byte pair encoding (BBPE) representations, and analyze their strengths and weaknesses. We focus on developing a single end-to-end model to support utterance-based bilingual ASR, where speakers do not alternate between two languages in a single utterance but may change languages across utterances. We conduct our experiments on English and Mandarin dictation tasks, and we find that BBPE with penalty schemes can improve utterance-based bilingual ASR performance by 2% to 5% relative even with smaller number of outputs and fewer parameters. We conclude with analysis that indicates directions for further improving multilingual ASR.

Keywords

Cite

@article{arxiv.2205.00485,
  title  = {Bilingual End-to-End ASR with Byte-Level Subwords},
  author = {Liuhui Deng and Roger Hsiao and Arnab Ghoshal},
  journal= {arXiv preprint arXiv:2205.00485},
  year   = {2022}
}

Comments

5 pages, to be published in IEEE ICASSP 2022

R2 v1 2026-06-24T11:03:56.095Z