English

Cross-lingual Extended Named Entity Classification of Wikipedia Articles

Computation and Language 2020-10-20 v2

Abstract

The FPT.AI team participated in the SHINRA2020-ML subtask of the NTCIR-15 SHINRA task. This paper describes our method to solving the problem and discusses the official results. Our method focuses on learning cross-lingual representations, both on the word level and document level for page classification. We propose a three-stage approach including multilingual model pre-training, monolingual model fine-tuning and cross-lingual voting. Our system is able to achieve the best scores for 25 out of 30 languages; and its accuracy gaps to the best performing systems of the other five languages are relatively small.

Keywords

Cite

@article{arxiv.2010.03424,
  title  = {Cross-lingual Extended Named Entity Classification of Wikipedia Articles},
  author = {The Viet Bui and Phuong Le-Hong},
  journal= {arXiv preprint arXiv:2010.03424},
  year   = {2020}
}

Comments

Accepted to NTCIR-15

R2 v1 2026-06-23T19:07:54.147Z