English

OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages

Computation and Language 2025-12-19 v3

Abstract

We present OpenNER 1.0, a standardized collection of openly-available named entity recognition (NER) datasets. OpenNER contains 36 NER corpora that span 52 languages, human-annotated in varying named entity ontologies. We correct annotation format issues, standardize the original datasets into a uniform representation with consistent entity type names across corpora, and provide the collection in a structure that enables research in multilingual and multi-ontology NER. We provide baseline results using three pretrained multilingual language models and two large language models to compare the performance of recent models and facilitate future research in NER. We find that no single model is best in all languages and that significant work remains to obtain high performance from LLMs on the NER task. OpenNER is released at https://github.com/bltlab/open-ner.

Keywords

Cite

@article{arxiv.2412.09587,
  title  = {OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages},
  author = {Chester Palen-Michel and Maxwell Pickering and Maya Kruse and Jonne Sälevä and Constantine Lignos},
  journal= {arXiv preprint arXiv:2412.09587},
  year   = {2025}
}

Comments

Published in the proceedings of EMNLP 2025

R2 v1 2026-06-28T20:32:59.150Z