English

CFMS: Towards Explainable and Fine-Grained Chinese Multimodal Sarcasm Detection Benchmark

Computation and Language 2026-04-21 v1 Artificial Intelligence

Abstract

Multimodal sarcasm detection has recently garnered significant attention. However, existing benchmarks suffer from coarse-grained annotations and limited cultural coverage, which hinder research into fine-grained semantic understanding. To address this, we construct CFMS, the first fine-grained multimodal sarcasm dataset tailored for Chinese social media. It comprises 2,796 high-quality image-text pairs and provides a triple-level annotation framework: sarcasm identification, target recognition, and explanation generation. We find that the fine-grained explanation annotations effectively guide AI in generating images with explicit sarcastic intent. Furthermore, we curate a high-consistency parallel Chinese-English metaphor subset (200 entries each), revealing significant limitations of current models in metaphoric reasoning. To overcome the constraints of traditional retrieval methods, we propose a Reinforcement Learning-augmented In-Context Learning strategy (PGDS) to dynamically optimize exemplar selection. Extensive experiments demonstrate that CFMS provides a solid foundation for building reliable multimodal sarcasm understanding systems, and the PGDS method significantly outperforms existing baselines on key tasks. Our data and code are available at https://anonymous.4open.science/r/CFMS-E8F9.

Keywords

Cite

@article{arxiv.2604.16372,
  title  = {CFMS: Towards Explainable and Fine-Grained Chinese Multimodal Sarcasm Detection Benchmark},
  author = {Junzhao Zhang and Hsiu-Yuan Huang and Chenming Tang and Yutong Yang and Yunfang Wu},
  journal= {arXiv preprint arXiv:2604.16372},
  year   = {2026}
}
R2 v1 2026-07-01T12:14:53.868Z