Unsupervised Keyword Extraction from Polish Legal Texts
Computation and Language
2014-11-04 v2
Abstract
In this work, we present an application of the recently proposed unsupervised keyword extraction algorithm RAKE to a corpus of Polish legal texts from the field of public procurement. RAKE is essentially a language and domain independent method. Its only language-specific input is a stoplist containing a set of non-content words. The performance of the method heavily depends on the choice of such a stoplist, which should be domain adopted. Therefore, we complement RAKE algorithm with an automatic approach to selecting non-content words, which is based on the statistical properties of term distribution.
Keywords
Cite
@article{arxiv.1408.3731,
title = {Unsupervised Keyword Extraction from Polish Legal Texts},
author = {Michał Jungiewicz and Michał Łopuszyński},
journal= {arXiv preprint arXiv:1408.3731},
year = {2014}
}