English

NorNE: Annotating Named Entities for Norwegian

Computation and Language 2020-03-09 v2

Abstract

This paper presents NorNE, a manually annotated corpus of named entities which extends the annotation of the existing Norwegian Dependency Treebank. Comprising both of the official standards of written Norwegian (Bokm{\aa}l and Nynorsk), the corpus contains around 600,000 tokens and annotates a rich set of entity types including persons, organizations, locations, geo-political entities, products, and events, in addition to a class corresponding to nominals derived from names. We here present details on the annotation effort, guidelines, inter-annotator agreement and an experimental analysis of the corpus using a neural sequence labeling architecture.

Cite

@article{arxiv.1911.12146,
  title  = {NorNE: Annotating Named Entities for Norwegian},
  author = {Fredrik Jørgensen and Tobias Aasmoe and Anne-Stine Ruud Husevåg and Lilja Øvrelid and Erik Velldal},
  journal= {arXiv preprint arXiv:1911.12146},
  year   = {2020}
}

Comments

Accepted for LREC 2020

R2 v1 2026-06-23T12:28:58.287Z