English

FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems

Computation and Language 2026-01-06 v1

Abstract

Hate speech (HS) is a critical issue in online discourse, and one promising strategy to counter it is through the use of counter-narratives (CNs). Datasets linking HS with CNs are essential for advancing counterspeech research. However, even flagship resources like CONAN (Chung et al., 2019) annotate only a sparse subset of all possible HS-CN pairs, limiting evaluation. We introduce FC-CONAN (Fully Connected CONAN), the first dataset created by exhaustively considering all combinations of 45 English HS messages and 129 CNs. A two-stage annotation process involving nine annotators and four validators produces four partitions-Diamond, Gold, Silver, and Bronze-that balance reliability and scale. None of the labeled pairs overlap with CONAN, uncovering hundreds of previously unlabelled positives. FC-CONAN enables more faithful evaluation of counterspeech retrieval systems and facilitates detailed error analysis. The dataset is publicly available.

Keywords

Cite

@article{arxiv.2601.01350,
  title  = {FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems},
  author = {Juan Junqueras and Florian Boudin and May-Myo Zin and Ha-Thanh Nguyen and Wachara Fungwacharakorn and Damián Ariel Furman and Akiko Aizawa and Ken Satoh},
  journal= {arXiv preprint arXiv:2601.01350},
  year   = {2026}
}

Comments

Presented at NeLaMKRR@KR, 2025 (arXiv:2511.09575)