English

The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes

Artificial Intelligence 2021-04-09 v3 Computation and Language Computer Vision and Pattern Recognition

Abstract

This work proposes a new challenge set for multimodal classification, focusing on detecting hate speech in multimodal memes. It is constructed such that unimodal models struggle and only multimodal models can succeed: difficult examples ("benign confounders") are added to the dataset to make it hard to rely on unimodal signals. The task requires subtle reasoning, yet is straightforward to evaluate as a binary classification problem. We provide baseline performance numbers for unimodal models, as well as for multimodal models with various degrees of sophistication. We find that state-of-the-art methods perform poorly compared to humans (64.73% vs. 84.7% accuracy), illustrating the difficulty of the task and highlighting the challenge that this important problem poses to the community.

Keywords

Cite

@article{arxiv.2005.04790,
  title  = {The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes},
  author = {Douwe Kiela and Hamed Firooz and Aravind Mohan and Vedanuj Goswami and Amanpreet Singh and Pratik Ringshia and Davide Testuggine},
  journal= {arXiv preprint arXiv:2005.04790},
  year   = {2021}
}

Comments

NeurIPS 2020

R2 v1 2026-06-23T15:26:30.554Z