Bilingual Character Representation for Efficiently Addressing Out-of-Vocabulary Words in Code-Switching Named Entity Recognition
Computation and Language
2019-06-11 v2
Abstract
We propose an LSTM-based model with hierarchical architecture on named entity recognition from code-switching Twitter data. Our model uses bilingual character representation and transfer learning to address out-of-vocabulary words. In order to mitigate data noise, we propose to use token replacement and normalization. In the 3rd Workshop on Computational Approaches to Linguistic Code-Switching Shared Task, we achieved second place with 62.76% harmonic mean F1-score for English-Spanish language pair without using any gazetteer and knowledge-based information.
Keywords
Cite
@article{arxiv.1805.12061,
title = {Bilingual Character Representation for Efficiently Addressing Out-of-Vocabulary Words in Code-Switching Named Entity Recognition},
author = {Genta Indra Winata and Chien-Sheng Wu and Andrea Madotto and Pascale Fung},
journal= {arXiv preprint arXiv:1805.12061},
year = {2019}
}
Comments
Accepted in "3rd Workshop in Computational Approaches in Linguistic Code-switching", ACL 2018