English

Low-resource speech recognition and dialect identification of Irish in a multi-task framework

Computation and Language 2024-05-03 v1 Artificial Intelligence Sound Audio and Speech Processing

Abstract

This paper explores the use of Hybrid CTC/Attention encoder-decoder models trained with Intermediate CTC (InterCTC) for Irish (Gaelic) low-resource speech recognition (ASR) and dialect identification (DID). Results are compared to the current best performing models trained for ASR (TDNN-HMM) and DID (ECAPA-TDNN). An optimal InterCTC setting is initially established using a Conformer encoder. This setting is then used to train a model with an E-branchformer encoder and the performance of both architectures are compared. A multi-task fine-tuning approach is adopted for language model (LM) shallow fusion. The experiments yielded an improvement in DID accuracy of 10.8% relative to a baseline ECAPA-TDNN, and WER performance approaching the TDNN-HMM model. This multi-task approach emerges as a promising strategy for Irish low-resource ASR and DID.

Keywords

Cite

@article{arxiv.2405.01293,
  title  = {Low-resource speech recognition and dialect identification of Irish in a multi-task framework},
  author = {Liam Lonergan and Mengjie Qian and Neasa Ní Chiaráin and Christer Gobl and Ailbhe Ní Chasaide},
  journal= {arXiv preprint arXiv:2405.01293},
  year   = {2024}
}

Comments

7 pages. Accepted to Odyssey 2024 - The Speaker and Language Recognition Workshop

R2 v1 2026-06-28T16:14:02.776Z