English

GaelEval: Benchmarking LLM Performance for Scottish Gaelic

Computation and Language 2026-04-03 v1

Abstract

Multilingual large language models (LLMs) often exhibit emergent 'shadow' capabilities in languages without official support, yet their performance on these languages remains uneven and under-measured. This is particularly acute for morphosyntactically rich minority languages such as Scottish Gaelic, where translation benchmarks fail to capture structural competence. We introduce GaelEval, the first multi-dimensional benchmark for Gaelic, comprising: (i) an expert-authored morphosyntactic MCQA task; (ii) a culturally grounded translation benchmark and (iii) a large-scale cultural knowledge Q&A task. Evaluating 19 LLMs against a fluent-speaker human baseline (n=30n=30), we find that Gemini 3 Pro Preview achieves 83.3%83.3\% accuracy on the linguistic task, surpassing the human baseline (78.1%78.1\%). Proprietary models consistently outperform open-weight systems, and in-language (Gaelic) prompting yields a small but stable advantage (+2.4%2.4\%). On the cultural task, leading models exceed 90%90\% accuracy, though most systems perform worse under Gaelic prompting and absolute scores are inflated relative to the manual benchmark. Overall, GaelEval reveals that frontier models achieve above-human performance on several dimensions of Gaelic grammar, demonstrates the effect of Gaelic prompting and shows a consistent performance gap favouring proprietary over open-weight models.

Keywords

Cite

@article{arxiv.2604.02135,
  title  = {GaelEval: Benchmarking LLM Performance for Scottish Gaelic},
  author = {Peter Devine and William Lamb and Beatrice Alex and Ignatius Ezeani and Dawn Knight and Mícheál J. Ó Meachair and Paul Rayson and Martin Wynne},
  journal= {arXiv preprint arXiv:2604.02135},
  year   = {2026}
}

Comments

13 pages, to be published in Proceedings of LLMs4SSH (workshop co-located with LREC 2026; Mallorca, Spain; May 2026)

R2 v1 2026-07-01T11:51:10.568Z