English

Semantic Matching Against a Corpus: New Applications and Methods

Computation and Language 2018-08-30 v1

Abstract

We consider the case of a domain expert who wishes to explore the extent to which a particular idea is expressed in a text collection. We propose the task of semantically matching the idea, expressed as a natural language proposition, against a corpus. We create two preliminary tasks derived from existing datasets, and then introduce a more realistic one on disaster recovery designed for emergency managers, whom we engaged in a user study. On the latter, we find that a new model built from natural language entailment data produces higher-quality matches than simple word-vector averaging, both on expert-crafted queries and on ones produced by the subjects themselves. This work provides a proof-of-concept for such applications of semantic matching and illustrates key challenges.

Keywords

Cite

@article{arxiv.1808.09502,
  title  = {Semantic Matching Against a Corpus: New Applications and Methods},
  author = {Lucy H. Lin and Scott Miles and Noah A. Smith},
  journal= {arXiv preprint arXiv:1808.09502},
  year   = {2018}
}

Comments

18 pages, 5 figures

R2 v1 2026-06-23T03:47:00.847Z