English

Towards dialect-inclusive recognition in a low-resource language: are balanced corpora the answer?

Computation and Language 2023-07-17 v1 Sound Audio and Speech Processing

Abstract

ASR systems are generally built for the spoken 'standard', and their performance declines for non-standard dialects/varieties. This is a problem for a language like Irish, where there is no single spoken standard, but rather three major dialects: Ulster (Ul), Connacht (Co) and Munster (Mu). As a diagnostic to quantify the effect of the speaker's dialect on recognition performance, 12 ASR systems were trained, firstly using baseline dialect-balanced training corpora, and then using modified versions of the baseline corpora, where dialect-specific materials were either subtracted or added. Results indicate that dialect-balanced corpora do not yield a similar performance across the dialects: the Ul dialect consistently underperforms, whereas Mu yields lowest WERs. There is a close relationship between Co and Mu dialects, but one that is not symmetrical. These results will guide future corpus collection and system building strategies to optimise for cross-dialect performance equity.

Keywords

Cite

@article{arxiv.2307.07295,
  title  = {Towards dialect-inclusive recognition in a low-resource language: are balanced corpora the answer?},
  author = {Liam Lonergan and Mengjie Qian and Neasa Ní Chiaráin and Christer Gobl and Ailbhe Ní Chasaide},
  journal= {arXiv preprint arXiv:2307.07295},
  year   = {2023}
}

Comments

Accepted to Interspeech 2023, Dublin

R2 v1 2026-06-28T11:30:24.489Z