English

Semantic Matching of Documents from Heterogeneous Collections: A Simple and Transparent Method for Practical Applications

Computation and Language 2019-04-30 v1

Abstract

We present a very simple, unsupervised method for the pairwise matching of documents from heterogeneous collections. We demonstrate our method with the Concept-Project matching task, which is a binary classification task involving pairs of documents from heterogeneous collections. Although our method only employs standard resources without any domain- or task-specific modifications, it clearly outperforms the more complex system of the original authors. In addition, our method is transparent, because it provides explicit information about how a similarity score was computed, and efficient, because it is based on the aggregation of (pre-computable) word-level similarities.

Keywords

Cite

@article{arxiv.1904.12550,
  title  = {Semantic Matching of Documents from Heterogeneous Collections: A Simple and Transparent Method for Practical Applications},
  author = {Mark-Christoph Müller},
  journal= {arXiv preprint arXiv:1904.12550},
  year   = {2019}
}

Comments

To appear at RELATIONS 2019 workshop