English

AlbNER: A Corpus for Named Entity Recognition in Albanian

Computation and Language 2023-09-19 v1 Artificial Intelligence Machine Learning

Abstract

Scarcity of resources such as annotated text corpora for under-resourced languages like Albanian is a serious impediment in computational linguistics and natural language processing research. This paper presents AlbNER, a corpus of 900 sentences with labeled named entities, collected from Albanian Wikipedia articles. Preliminary results with BERT and RoBERTa variants fine-tuned and tested with AlbNER data indicate that model size has slight impact on NER performance, whereas language transfer has a significant one. AlbNER corpus and these obtained results should serve as baselines for future experiments.

Keywords

Cite

@article{arxiv.2309.08741,
  title  = {AlbNER: A Corpus for Named Entity Recognition in Albanian},
  author = {Erion Çano},
  journal= {arXiv preprint arXiv:2309.08741},
  year   = {2023}
}

Comments

5 pages, 6 tables

R2 v1 2026-06-28T12:23:07.680Z