Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum
Abstract
There is increasing interest in looking at dialects in NLP. However, most work to date still treats dialects as discrete categories. For instance, evaluative work in variation-oriented NLP for English often works with Indian English or African-American Venacular English as homogeneous categories (Faisal et al., 2024; Ziems et al., 2023), yet even within one variety there is substantial variation. We examine within-dialect variation and show that performance critically varies within categories. We measure speech-to-text performance on Italian dialects, and empirically observe a geographical performance disparity. This disparity correlates substantially (-0.5) with linguistic similarity to the highest performing dialect variety. We cross-examine our results against dialectometry methods, and interpret the performance disparity to be due to a bias towards dialects that are more similar to the standard variety in the speech-to-text model examined. We additionally leverage geostatistical methods to predict zero-shot performance at unseen sites, and find the incorporation of geographical information to substantially improve prediction performance, indicating there to be geographical structure in the performance distribution.
Keywords
Cite
@article{arxiv.2410.14589,
title = {Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum},
author = {Ryan Soh-Eun Shim and Barbara Plank},
journal= {arXiv preprint arXiv:2410.14589},
year = {2025}
}
Comments
Published in NAACL 2025 findings