English

Novel Keyword Extraction and Language Detection Approaches

Computation and Language 2020-09-25 v1

Abstract

Fuzzy string matching and language classification are important tools in Natural Language Processing pipelines, this paper provides advances in both areas. We propose a fast novel approach to string tokenisation for fuzzy language matching and experimentally demonstrate an 83.6% decrease in processing time with an estimated improvement in recall of 3.1% at the cost of a 2.6% decrease in precision. This approach is able to work even where keywords are subdivided into multiple words, without needing to scan character-to-character. So far there has been little work considering using metadata to enhance language classification algorithms. We provide observational data and find the Accept-Language header is 14% more likely to match the classification than the IP Address.

Keywords

Cite

@article{arxiv.2009.11832,
  title  = {Novel Keyword Extraction and Language Detection Approaches},
  author = {Malgorzata Pikies and Andronicus Riyono and Junade Ali},
  journal= {arXiv preprint arXiv:2009.11832},
  year   = {2020}
}
R2 v1 2026-06-23T18:46:30.657Z