English

A Multi-way Parallel Named Entity Annotated Corpus for English, Tamil and Sinhala

Computation and Language 2025-01-16 v2

Abstract

This paper presents a multi-way parallel English-Tamil-Sinhala corpus annotated with Named Entities (NEs), where Sinhala and Tamil are low-resource languages. Using pre-trained multilingual Language Models (mLMs), we establish new benchmark Named Entity Recognition (NER) results on this dataset for Sinhala and Tamil. We also carry out a detailed investigation on the NER capabilities of different types of mLMs. Finally, we demonstrate the utility of our NER system on a low-resource Neural Machine Translation (NMT) task. Our dataset is publicly released: https://github.com/suralk/multiNER.

Keywords

Cite

@article{arxiv.2412.02056,
  title  = {A Multi-way Parallel Named Entity Annotated Corpus for English, Tamil and Sinhala},
  author = {Surangika Ranathunga and Asanka Ranasinghea and Janaka Shamala and Ayodya Dandeniyaa and Rashmi Galappaththia and Malithi Samaraweeraa},
  journal= {arXiv preprint arXiv:2412.02056},
  year   = {2025}
}
R2 v1 2026-06-28T20:20:38.379Z