English
Related papers

Related papers: UnsafeBench: Benchmarking Image Safety Classifiers…

200 papers

Recent advances in multi-modal generative models have driven substantial improvements in image editing. However, current generative models still struggle with handling diverse and complex image editing tasks that require implicit reasoning,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Feng Han , Yibin Wang , Chenglin Li , Zheming Liang , Dianyi Wang , Yang Jiao , Zhipeng Wei , Chao Gong , Cheng Jin , Jingjing Chen , Jiaqi Wang

Image content safety has become a significant challenge with the rise of visual media on online platforms. Meanwhile, in the age of AI-generated content (AIGC), many image generation models are capable of producing harmful content, such as…

One of the key challenges of detecting AI-generated images is spotting images that have been created by previously unseen generative models. We argue that the limited diversity of the training data is a major obstacle to addressing this…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Jeongsoo Park , Andrew Owens

The threats posed by AI-generated media, particularly deepfakes, are now raising significant challenges for multimedia forensics, misinformation detection, and biometric system resulting in erosion of public trust in the legal system,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Nusrat Tasnim , Kutub Uddin , Khalid Mahmood Malik

We introduce a novel framework for AI-generated image detection through epistemic uncertainty, aiming to address critical security concerns in the era of generative models. Our key insight stems from the observation that distributional…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Jun Nie , Yonggang Zhang , Tongliang Liu , Yiu-ming Cheung , Bo Han , Xinmei Tian

Text-to-image models are increasingly popular and impactful, yet concerns regarding their safety and fairness remain. This study investigates the ability of ten popular Stable Diffusion models to generate harmful images, including NSFW,…

Computers and Society · Computer Science 2025-08-29 Matthias Schneider , Thilo Hagendorff

Studying the robustness of Large Language Models (LLMs) to unsafe behaviors is an important topic of research today. Building safety classification models or guard models, which are fine-tuned models for input/output safety classification…

Computation and Language · Computer Science 2025-07-30 Sowmya Vajjala

As a research community, we are still lacking a systematic understanding of the progress on adversarial robustness which often makes it hard to identify the most promising ideas in training robust models. A key challenge in benchmarking…

Despite the promising performance of existing visual models on public benchmarks, the critical assessment of their robustness for real-world applications remains an ongoing challenge. To bridge this gap, we propose an explainable visual…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Qiang Li , Dan Zhang , Shengzhao Lei , Xun Zhao , Porawit Kamnoedboon , WeiWei Li , Junhao Dong , Shuyan Li

Large-scale image datasets frequently contain identifiable or sensitive content, raising privacy risks when training models that may memorize and leak such information. We present Unsafe2Safe, a fully automated pipeline that detects…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Mih Dinh , SouYoung Jin

AI Safety Moderation (ASM) classifiers are designed to moderate content on social media platforms and to serve as guardrails that prevent Large Language Models (LLMs) from being fine-tuned on unsafe inputs. Owing to their potential for…

Computation and Language · Computer Science 2025-01-24 Akshit Achara , Anshuman Chhabra

AI-generated images have reached a quality level at which humans are incapable of reliably distinguishing them from real images. To counteract the inherent risk of fraud and disinformation, the detection of AI-generated images is a pressing…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Hicham Eddoubi , Jonas Ricker , Federico Cocchi , Lorenzo Baraldi , Angelo Sotgiu , Maura Pintor , Marcella Cornia , Lorenzo Baraldi , Asja Fischer , Rita Cucchiara , Battista Biggio

The rapid advancement of vision-language models (VLMs) has brought a lot of attention to their safety alignment. However, existing methods have primarily focused on model undersafety, where the model responds to hazardous queries, while…

Information Retrieval · Computer Science 2025-05-28 Jiahui Geng , Qing Li , Zongxiong Chen , Yuxia Wang , Derui Zhu , Zhuohan Xie , Chenyang Lyu , Xiuying Chen , Preslav Nakov , Fakhri Karray

The robust safety of Vision-Language Large Models (VLLMs) against joint multilingual and multimodal threats remains severely underexplored. Current benchmarks typically isolate these dimensions, being either multilingual but text-only, or…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Enyi Shi , Pengyang Shao , Yanxin Zhang , Chenhang Cui , Jiayi Lyu , Xiaobo Xia , Fei Shen , Tat-Seng Chua

The rapid progress of Generative Artificial Intelligence (GenAI) has enabled the effortless synthesis of high-quality visual content, while simultaneously raising pressing concerns about intellectual property protection, authenticity, and…

Cryptography and Security · Computer Science 2026-03-17 Jie Cao , Qi Li , Zelin Zhang , Jianbing Ni , Rongxing Lu

Text-to-image (T2I) models, such as Stable Diffusion, have exhibited remarkable performance in generating high-quality images from text descriptions in recent years. However, text-to-image models may be tricked into generating…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Xinfeng Li , Yuchen Yang , Jiangyi Deng , Chen Yan , Yanjiao Chen , Xiaoyu Ji , Wenyuan Xu

Rapid deployment of vision-language models (VLMs) magnifies safety risks, yet most evaluations rely on artificial images. This study asks: How safe are current VLMs when confronted with meme images that ordinary users share? To investigate…

Computation and Language · Computer Science 2025-09-24 DongGeon Lee , Joonwon Jang , Jihae Jeong , Hwanjo Yu

The ability to distinguish whether an image is generated by artificial intelligence (AI) is a crucial ingredient in human intelligence, usually accompanied by a complex and dialectical forensic and reasoning process. However, current fake…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Yixuan Li , Xuelin Liu , Xiaoyang Wang , Bu Sung Lee , Shiqi Wang , Anderson Rocha , Weisi Lin

Automated content filtering and moderation is an important tool that allows online platforms to build striving user communities that facilitate cooperation and prevent abuse. Unfortunately, resourceful actors try to bypass automated filters…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Florian Stimberg , Ayan Chakrabarti , Chun-Ta Lu , Hussein Hazimeh , Otilia Stretcu , Wei Qiao , Yintao Liu , Merve Kaya , Cyrus Rashtchian , Ariel Fuxman , Mehmet Tek , Sven Gowal

An important challenge when using computer vision models in the real world is to evaluate their performance in potential out-of-distribution (OOD) scenarios. While simple synthetic corruptions are commonly applied to test OOD robustness,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Olaf Dünkel , Artur Jesslen , Jiahao Xie , Christian Theobalt , Christian Rupprecht , Adam Kortylewski