English

A Joint Approach to Compound Splitting and Idiomatic Compound Detection

Computation and Language 2020-03-24 v1

Abstract

Applications such as machine translation, speech recognition, and information retrieval require efficient handling of noun compounds as they are one of the possible sources for out-of-vocabulary (OOV) words. In-depth processing of noun compounds requires not only splitting them into smaller components (or even roots) but also the identification of instances that should remain unsplitted as they are of idiomatic nature. We develop a two-fold deep learning-based approach of noun compound splitting and idiomatic compound detection for the German language that we train using a newly collected corpus of annotated German compounds. Our neural noun compound splitter operates on a sub-word level and outperforms the current state of the art by about 5%.

Keywords

Cite

@article{arxiv.2003.09606,
  title  = {A Joint Approach to Compound Splitting and Idiomatic Compound Detection},
  author = {Irina Krotova and Sergey Aksenov and Ekaterina Artemova},
  journal= {arXiv preprint arXiv:2003.09606},
  year   = {2020}
}

Comments

8 pages, 5 tables, 1 figure, accepted at LREC 2020

R2 v1 2026-06-23T14:22:21.924Z