English

Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding

Computation and Language 2026-04-21 v4

Abstract

Negation is a fundamental linguistic phenomenon that poses ongoing challenges for Large Language Models (LLMs), particularly in tasks requiring deep semantic understanding. Current benchmarks often treat negation as a minor detail within broader tasks, such as natural language inference. Consequently, there is a lack of benchmarks specifically designed to evaluate comprehension of negation. In this work, we introduce Thunder-NUBench, a novel benchmark explicitly created to assess sentence-level understanding of negation in LLMs. Thunder-NUBench goes beyond merely identifying surface-level cues by contrasting standard negation with structurally diverse alternatives, such as local negation, contradiction, and paraphrase. This benchmark includes manually curated sentence-negation pairs and a multiple-choice dataset, allowing for a comprehensive evaluation of models' understanding of negation.

Keywords

Cite

@article{arxiv.2506.14397,
  title  = {Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding},
  author = {Yeonkyoung So and Gyuseong Lee and Sungmok Jung and Joonhak Lee and JiA Kang and Sangho Kim and Jaejin Lee},
  journal= {arXiv preprint arXiv:2506.14397},
  year   = {2026}
}
R2 v1 2026-07-01T03:21:38.901Z