English

Bilingual Character Representation for Efficiently Addressing Out-of-Vocabulary Words in Code-Switching Named Entity Recognition

Computation and Language 2019-06-11 v2

Abstract

We propose an LSTM-based model with hierarchical architecture on named entity recognition from code-switching Twitter data. Our model uses bilingual character representation and transfer learning to address out-of-vocabulary words. In order to mitigate data noise, we propose to use token replacement and normalization. In the 3rd Workshop on Computational Approaches to Linguistic Code-Switching Shared Task, we achieved second place with 62.76% harmonic mean F1-score for English-Spanish language pair without using any gazetteer and knowledge-based information.

Keywords

Cite

@article{arxiv.1805.12061,
  title  = {Bilingual Character Representation for Efficiently Addressing Out-of-Vocabulary Words in Code-Switching Named Entity Recognition},
  author = {Genta Indra Winata and Chien-Sheng Wu and Andrea Madotto and Pascale Fung},
  journal= {arXiv preprint arXiv:1805.12061},
  year   = {2019}
}

Comments

Accepted in "3rd Workshop in Computational Approaches in Linguistic Code-switching", ACL 2018