English

Reliable Evaluations for Natural Language Inference based on a Unified Cross-dataset Benchmark

Computation and Language 2020-10-16 v1 Artificial Intelligence

Abstract

Recent studies show that crowd-sourced Natural Language Inference (NLI) datasets may suffer from significant biases like annotation artifacts. Models utilizing these superficial clues gain mirage advantages on the in-domain testing set, which makes the evaluation results over-estimated. The lack of trustworthy evaluation settings and benchmarks stalls the progress of NLI research. In this paper, we propose to assess a model's trustworthy generalization performance with cross-datasets evaluation. We present a new unified cross-datasets benchmark with 14 NLI datasets, and re-evaluate 9 widely-used neural network-based NLI models as well as 5 recently proposed debiasing methods for annotation artifacts. Our proposed evaluation scheme and experimental baselines could provide a basis to inspire future reliable NLI research.

Keywords

Cite

@article{arxiv.2010.07676,
  title  = {Reliable Evaluations for Natural Language Inference based on a Unified Cross-dataset Benchmark},
  author = {Guanhua Zhang and Bing Bai and Jian Liang and Kun Bai and Conghui Zhu and Tiejun Zhao},
  journal= {arXiv preprint arXiv:2010.07676},
  year   = {2020}
}
R2 v1 2026-06-23T19:22:20.999Z