English

ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs

Computation and Language 2026-03-30 v1 Artificial Intelligence Machine Learning

Abstract

As Large Language Models (LLMs) expand across multilingual domains, evaluating their performance in under-represented languages becomes increasingly important. European Portuguese (pt-PT) is particularly affected, as existing training data and benchmarks are mainly in Brazilian Portuguese (pt-BR). To address this, we introduce ALBA, a linguistically grounded benchmark designed from the ground up to assess LLM proficiency in linguistic-related tasks in pt-PT across eight linguistic dimensions, including Language Variety, Culture-bound Semantics, Discourse Analysis, Word Plays, Syntax, Morphology, Lexicology, and Phonetics and Phonology. ALBA is manually constructed by language experts and paired with an LLM-as-a-judge framework for scalable evaluation of pt-PT generated language. Experiments on a diverse set of models reveal performance variability across linguistic dimensions, highlighting the need for comprehensive, variety-sensitive benchmarks that support further development of tools in pt-PT.

Keywords

Cite

@article{arxiv.2603.26516,
  title  = {ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs},
  author = {Inês Vieira and Inês Calvo and Iago Paulo and James Furtado and Rafael Ferreira and Diogo Tavares and Diogo Glória-Silva and David Semedo and João Magalhães},
  journal= {arXiv preprint arXiv:2603.26516},
  year   = {2026}
}

Comments

PROPOR 2026 - The 17th International Conference on Computational Processing of Portuguese