从反转 CLIP 模型中学到了什么?
计算机视觉与模式识别
2024-03-06 v1 机器学习
摘要
我们采用基于反转的方法来检验 CLIP 模型。我们的研究表明,反转 CLIP 模型会生成与指定目标提示在语义上对齐的图像。我们利用这些反转图像来洞察 CLIP 模型的各个方面,例如其概念混合能力以及性别偏见的包含情况。我们特别观察到在模型反转过程中出现了 NSFW(Not Safe For Work,不适合工作场所)图像。这种现象甚至发生在语义上无害的提示(如“美丽的风景”)以及涉及名人姓名的提示中。
引用
@article{arxiv.2403.02580,
title = {What do we learn from inverting CLIP models?},
author = {Hamid Kazemi and Atoosa Chegini and Jonas Geiping and Soheil Feizi and Tom Goldstein},
journal= {arXiv preprint arXiv:2403.02580},
year = {2024}
}
备注
Warning: This paper contains sexually explicit images and language, offensive visuals and terminology, discussions on pornography, gender bias, and other potentially unsettling, distressing, and/or offensive content for certain readers