English

Disturbing Image Detection Using LMM-Elicited Emotion Embeddings

Computer Vision and Pattern Recognition 2024-06-19 v1

Abstract

In this paper we deal with the task of Disturbing Image Detection (DID), exploiting knowledge encoded in Large Multimodal Models (LMMs). Specifically, we propose to exploit LMM knowledge in a two-fold manner: first by extracting generic semantic descriptions, and second by extracting elicited emotions. Subsequently, we use the CLIP's text encoder in order to obtain the text embeddings of both the generic semantic descriptions and LMM-elicited emotions. Finally, we use the aforementioned text embeddings along with the corresponding CLIP's image embeddings for performing the DID task. The proposed method significantly improves the baseline classification accuracy, achieving state-of-the-art performance on the augmented Disturbing Image Detection dataset.

Keywords

Cite

@article{arxiv.2406.12668,
  title  = {Disturbing Image Detection Using LMM-Elicited Emotion Embeddings},
  author = {Maria Tzelepi and Vasileios Mezaris},
  journal= {arXiv preprint arXiv:2406.12668},
  year   = {2024}
}

Comments

Accepted for publication, LVLM Workshop @ IEEE Int. Conf. on Image Processing (ICIP 2024), Abu Dhabi, United Arab Emirates, Oct. 2024. This is the authors' "accepted version"

R2 v1 2026-06-28T17:10:28.418Z