English

Probing Commonsense Reasoning Capability of Text-to-Image Generative Models via Non-visual Description

Multimedia 2024-01-24 v2

Abstract

Commonsense reasoning, the ability to make logical assumptions about daily scenes, is one core intelligence of human beings. In this work, we present a novel task and dataset for evaluating the ability of text-to-image generative models to conduct commonsense reasoning, which we call PAINTaboo. Given a description with few visual clues of one object, the goal is to generate images illustrating the object correctly. The dataset was carefully hand-curated and covered diverse object categories to analyze model performance comprehensively. Our investigation of several prevalent text-to-image generative models reveals that these models are not proficient in commonsense reasoning, as anticipated. We trust that PAINTaboo can improve our understanding of the reasoning abilities of text-to-image generative models.

Keywords

Cite

@article{arxiv.2312.07294,
  title  = {Probing Commonsense Reasoning Capability of Text-to-Image Generative Models via Non-visual Description},
  author = {Mianzhi Pan and Jianfei Li and Mingyue Yu and Zheng Ma and Kanzhi Cheng and Jianbing Zhang and Jiajun Chen},
  journal= {arXiv preprint arXiv:2312.07294},
  year   = {2024}
}

Comments

It is an incomplete work