English

Parameter-efficient Adaptation of Multilingual Multimodal Models for Low-resource ASR

Computation and Language 2024-10-18 v1 Artificial Intelligence Machine Learning Audio and Speech Processing

Abstract

Automatic speech recognition (ASR) for low-resource languages remains a challenge due to the scarcity of labeled training data. Parameter-efficient fine-tuning and text-only adaptation are two popular methods that have been used to address such low-resource settings. In this work, we investigate how these techniques can be effectively combined using a multilingual multimodal model like SeamlessM4T. Multimodal models are able to leverage unlabeled text via text-only adaptation with further parameter-efficient ASR fine-tuning, thus boosting ASR performance. We also show cross-lingual transfer from a high-resource language, achieving up to a relative 17% WER reduction over a baseline in a zero-shot setting without any labeled speech.

Keywords

Cite

@article{arxiv.2410.13445,
  title  = {Parameter-efficient Adaptation of Multilingual Multimodal Models for Low-resource ASR},
  author = {Abhishek Gupta and Amruta Parulekar and Sameep Chattopadhyay and Preethi Jyothi},
  journal= {arXiv preprint arXiv:2410.13445},
  year   = {2024}
}
R2 v1 2026-06-28T19:25:41.223Z