English

CLASH: A Benchmark for Cross-Modal Contradiction Detection

Computer Vision and Pattern Recognition 2025-11-25 v1 Artificial Intelligence Machine Learning

Abstract

Contradictory multimodal inputs are common in real-world settings, yet existing benchmarks typically assume input consistency and fail to evaluate cross-modal contradiction detection - a fundamental capability for preventing hallucinations and ensuring reliability. We introduce CLASH, a novel benchmark for multimodal contradiction detection, featuring COCO images paired with contradictory captions containing controlled object-level or attribute-level contradictions. The samples include targeted questions evaluated in both multiple-choice and open-ended formats. The benchmark provides an extensive fine-tuning set filtered through automated quality checks, alongside a smaller human-verified diagnostic set. Our analysis of state-of-the-art models reveals substantial limitations in recognizing cross-modal conflicts, exposing systematic modality biases and category-specific weaknesses. Furthermore, we empirically demonstrate that targeted fine-tuning on CLASH substantially enhances conflict detection capabilities.

Keywords

Cite

@article{arxiv.2511.19199,
  title  = {CLASH: A Benchmark for Cross-Modal Contradiction Detection},
  author = {Teodora Popordanoska and Jiameng Li and Matthew B. Blaschko},
  journal= {arXiv preprint arXiv:2511.19199},
  year   = {2025}
}

Comments

First two authors contributed equally

R2 v1 2026-07-01T07:52:18.167Z