English

Analysis of Named Entity Recognition and Linking for Tweets

Computation and Language 2014-11-26 v1

Abstract

Applying natural language processing for mining and intelligent information access to tweets (a form of microblog) is a challenging, emerging research area. Unlike carefully authored news text and other longer content, tweets pose a number of new challenges, due to their short, noisy, context-dependent, and dynamic nature. Information extraction from tweets is typically performed in a pipeline, comprising consecutive stages of language identification, tokenisation, part-of-speech tagging, named entity recognition and entity disambiguation (e.g. with respect to DBpedia). In this work, we describe a new Twitter entity disambiguation dataset, and conduct an empirical analysis of named entity recognition and disambiguation, investigating how robust a number of state-of-the-art systems are on such noisy texts, what the main sources of error are, and which problems should be further investigated to improve the state of the art.

Keywords

Cite

@article{arxiv.1410.7182,
  title  = {Analysis of Named Entity Recognition and Linking for Tweets},
  author = {Leon Derczynski and Diana Maynard and Giuseppe Rizzo and Marieke van Erp and Genevieve Gorrell and Raphaël Troncy and Johann Petrak and Kalina Bontcheva},
  journal= {arXiv preprint arXiv:1410.7182},
  year   = {2014}
}

Comments

35 pages, accepted to journal Information Processing and Management

R2 v1 2026-06-22T06:37:10.394Z