English

Collecting Diverse Natural Language Inference Problems for Sentence Representation Evaluation

Computation and Language 2018-08-30 v2

Abstract

We present a large-scale collection of diverse natural language inference (NLI) datasets that help provide insight into how well a sentence representation captures distinct types of reasoning. The collection results from recasting 13 existing datasets from 7 semantic phenomena into a common NLI structure, resulting in over half a million labeled context-hypothesis pairs in total. We refer to our collection as the DNC: Diverse Natural Language Inference Collection. The DNC is available online at https://www.decomp.net, and will grow over time as additional resources are recast and added from novel sources.

Keywords

Cite

@article{arxiv.1804.08207,
  title  = {Collecting Diverse Natural Language Inference Problems for Sentence Representation Evaluation},
  author = {Adam Poliak and Aparajita Haldar and Rachel Rudinger and J. Edward Hu and Ellie Pavlick and Aaron Steven White and Benjamin Van Durme},
  journal= {arXiv preprint arXiv:1804.08207},
  year   = {2018}
}

Comments

To be presented at EMNLP 2018. 15 pages

R2 v1 2026-06-23T01:31:56.686Z