Yor-Sarc: A gold-standard dataset for sarcasm detection in a low-resource African language
Abstract
Sarcasm detection poses a fundamental challenge in computational semantics, requiring models to resolve disparities between literal and intended meaning. The challenge is amplified in low-resource languages where annotated datasets are scarce or nonexistent. We present \textbf{Yor-Sarc}, the first gold-standard dataset for sarcasm detection in Yor\`{u}b\'{a}, a tonal Niger-Congo language spoken by over million people. The dataset comprises 436 instances annotated by three native speakers from diverse dialectal backgrounds using an annotation protocol specifically designed for Yor\`{u}b\'{a} sarcasm by taking culture into account. This protocol incorporates context-sensitive interpretation and community-informed guidelines and is accompanied by a comprehensive analysis of inter-annotator agreement to support replication in other African languages. Substantial to almost perfect agreement was achieved (Fleiss' ; pairwise Cohen's --), with unanimous consensus. One annotator pair achieved almost perfect agreement (; raw agreement), exceeding a number of reported benchmarks for English sarcasm research works. The remaining majority-agreement cases are preserved as soft labels for uncertainty-aware modelling. Yor-Sarc\footnote{https://github.com/toheebadura/yor-sarc} is expected to facilitate research on semantic interpretation and culturally informed NLP for low-resource African languages.
Keywords
Cite
@article{arxiv.2602.18964,
title = {Yor-Sarc: A gold-standard dataset for sarcasm detection in a low-resource African language},
author = {Toheeb Aduramomi Jimoh and Tabea De Wille and Nikola S. Nikolov},
journal= {arXiv preprint arXiv:2602.18964},
year = {2026}
}