English

An Improved Method for Class-specific Keyword Extraction: A Case Study in the German Business Registry

Computation and Language 2024-07-22 v1

Abstract

The task of keyword extraction\textit{keyword extraction} is often an important initial step in unsupervised information extraction, forming the basis for tasks such as topic modeling or document classification. While recent methods have proven to be quite effective in the extraction of keywords, the identification of class-specific\textit{class-specific} keywords, or only those pertaining to a predefined class, remains challenging. In this work, we propose an improved method for class-specific keyword extraction, which builds upon the popular KeyBERT\textbf{KeyBERT} library to identify only keywords related to a class described by seed keywords\textit{seed keywords}. We test this method using a dataset of German business registry entries, where the goal is to classify each business according to an economic sector. Our results reveal that our method greatly improves upon previous approaches, setting a new standard for class-specific\textit{class-specific} keyword extraction.

Keywords

Cite

@article{arxiv.2407.14085,
  title  = {An Improved Method for Class-specific Keyword Extraction: A Case Study in the German Business Registry},
  author = {Stephen Meisenbacher and Tim Schopf and Weixin Yan and Patrick Holl and Florian Matthes},
  journal= {arXiv preprint arXiv:2407.14085},
  year   = {2024}
}

Comments

7 pages, 1 figure, 1 table. Accepted to KONVENS 2024

R2 v1 2026-06-28T17:46:57.895Z