English

Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling

Audio and Speech Processing 2024-12-24 v1 Computation and Language

Abstract

Multilingual Automatic Speech Recognition (ASR) aims to recognize and transcribe speech from multiple languages within a single system. Whisper, one of the most advanced ASR models, excels in this domain by handling 99 languages effectively, leveraging a vast amount of data and incorporating language tags as prefixes to guide the recognition process. However, despite its success, Whisper struggles with unseen languages, those not included in its pre-training. Motivated by the observation that many languages share linguistic characteristics, we propose methods that exploit these relationships to enhance ASR performance on unseen languages. Specifically, we introduce a weighted sum method, which computes a weighted sum of the embeddings of language tags, using Whisper's predicted language probabilities. In addition, we develop a predictor-based approach that refines the weighted sum embedding to more closely approximate the true embedding for unseen languages. Experimental results demonstrate substantial improvements in ASR performance, both in zero-shot and fine-tuning settings. Our proposed methods outperform baseline approaches, providing an effective solution for addressing unseen languages in multilingual ASR.

Keywords

Cite

@article{arxiv.2412.16474,
  title  = {Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling},
  author = {Shao-Syuan Huang and Kuan-Po Huang and Andy T. Liu and Hung-yi Lee},
  journal= {arXiv preprint arXiv:2412.16474},
  year   = {2024}
}

Comments

Accepted by ICASSP 2025

R2 v1 2026-06-28T20:44:42.421Z