English
Related papers

Related papers: The Violation State: Safety State Persistence in a…

200 papers

The proliferation of generative AI has transformed creative workflows, yet current systems face critical challenges in controllability and content protection. We propose a novel multi-agent framework that addresses both limitations through…

Multiagent Systems · Computer Science 2026-01-21 Haris Khan , Sadia Asif

Continuous authentication systems leveraging free-text keyboard dynamics offer a promising additional layer of security in a multifactor authentication setup that can be used in a transparent way with no impact on user experience. This…

Machine Learning · Computer Science 2025-05-09 Roberto Dillon , Arushi

Scams exploiting real-time social engineering -- such as phishing, impersonation, and phone fraud -- remain a persistent and evolving threat across digital platforms. Existing defenses are largely reactive, offering limited protection…

Cryptography and Security · Computer Science 2026-01-21 Ismail Hossain , Sai Puppala , Md Jahangir Alam , Sajedul Talukder

Moral reasoning is fundamental to safe Artificial Intelligence (AI), yet ensuring its consistency across modalities becomes critical as AI systems evolve from text-based assistants to embodied agents. Current safety techniques demonstrate…

Artificial Intelligence · Computer Science 2026-03-18 Xinyi Yang , Chenheng Xu , Weijun Hong , Ce Mo , Qian Wang , Fang Fang , Yixin Zhu

Large-scale vision-language models, especially CLIP, have demonstrated remarkable performance across diverse downstream tasks. Soft prompts, as carefully crafted modules that efficiently adapt vision-language models to specific tasks,…

Cryptography and Security · Computer Science 2026-05-27 Wenyuan Yang , Yichen Sun , Changzheng Chen , Zhixuan Chu , Jiaheng Zhang , Yiming Li , Dacheng Tao

Recently, stable diffusion (SD) models have typically flourished in the field of image synthesis and personalized editing, with a range of photorealistic and unprecedented images being successfully generated. As a result, widespread…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Zhiyuan Ma , Guoli Jia , Biqing Qi , Bowen Zhou

The use of large language models to assess user states in conversational and adaptive systems is based on the assumption that the metrics used for such assessment are stable and interpretable at the level of individual scores. This paper…

Artificial Intelligence · Computer Science 2026-05-18 Izabella Krzeminska , Michal Butkiewicz , Ewa Komkowska

Text-to-image (T2I) models have been widely applied in generating high-fidelity images across various domains. However, these models may also be abused to produce Not-Safe-for-Work (NSFW) content via jailbreak attacks. Existing jailbreak…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Xingkai Peng , Jun Jiang , Meng Tong , Shuai Li , Weiming Zhang , Nenghai Yu , Kejiang Chen

As Vision-Language Models (VLMs) move into interactive, multi-turn use, safety concerns intensify for multimodal multi-turn dialogue, which is characterized by concealment of malicious intent, contextual risk accumulation, and cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Guolei Huang , Qinzhi Peng , Gan Xu , Yao Huang , Yuxuan Lu , Yongjun Shen

This study explores ChatGPT's capabilities, stability, and risks in simulating psychological counseling sessions in a school counseling context. Using scripted role-plays between a human counselor and an AI client, we examine how a large…

Human-Computer Interaction · Computer Science 2025-12-08 Yang Ni , Yanzhuo Cao

Watermarking is becoming the default mechanism for AI content authentication, with governance policies and frameworks referencing it as infrastructure for content provenance. Yet across text, image, and audio modalities, watermark signal…

Computers and Society · Computer Science 2026-04-16 Alexander Nemecek , Osama Zafar , Yuqiao Xu , Wenbiao Li , Erman Ayday

When AI interacts with the physical world -- as a robot or an assistive agent -- new safety challenges emerge beyond those of purely ``digital AI". In such interactions, the potential for physical harm is direct and immediate. How well do…

The key problem facing ChatGPT-like AI's use across society is that its behavior can shift, unnoticed, from desirable to undesirable -- encouraging self-harm, extremist acts, financial losses, or costly medical and military mistakes -- and…

Artificial Intelligence · Computer Science 2026-05-15 Neil F. Johnson , Frank Yingjie Huo

Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly understood. Existing systems already span multiple process designs, including direct response…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Yuan Tian , Bing Hu , Fang Wu , Xiaomin Li , Binghang Lu , Neil Zhenqiang Gong

Dialogue safety remains a pervasive challenge in open-domain human-machine interaction. Existing approaches propose distinctive dialogue safety taxonomies and datasets for detecting explicitly harmful responses. However, these taxonomies…

Computation and Language · Computer Science 2023-08-01 Huachuan Qiu , Tong Zhao , Anqi Li , Shuai Zhang , Hongliang He , Zhenzhong Lan

This paper analyzes the community safety guidelines of five text-to-image (T2I) generation platforms and audits five T2I models, focusing on prompts related to the representation of humans in areas that might lead to societal stigma. While…

Computers and Society · Computer Science 2024-09-27 Piera Riccio , Georgina Curto , Nuria Oliver

AI-generated content is becoming increasingly prevalent in the real world, leading to serious ethical and societal concerns. For instance, adversaries might exploit large multimodal models (LMMs) to create images that violate ethical or…

Computation and Language · Computer Science 2025-04-14 Hongchao Fang , Yixin Liu , Jiangshu Du , Can Qin , Ran Xu , Feng Liu , Lichao Sun , Dongwon Lee , Lifu Huang , Wenpeng Yin

Large-scale pre-trained generative models are taking the world by storm, due to their abilities in generating creative content. Meanwhile, safeguards for these generative models are developed, to protect users' rights and safety, most of…

Cryptography and Security · Computer Science 2024-10-14 Guanlin Li , Kangjie Chen , Shudong Zhang , Jie Zhang , Tianwei Zhang

NSFW (Not Safe for Work) content, in the context of a dialogue, can have severe side effects on users in open-domain dialogue systems. However, research on detecting NSFW language, especially sexually explicit content, within a dialogue…

Computation and Language · Computer Science 2024-03-22 Huachuan Qiu , Shuai Zhang , Hongliang He , Anqi Li , Zhenzhong Lan

Generative models have rapidly evolved to generate realistic outputs. However, their synthetic outputs increasingly challenge the clear distinction between natural and AI-generated content, necessitating robust watermarking techniques.…

Machine Learning · Computer Science 2026-05-20 Kasra Arabi , R. Teal Witter , Chinmay Hegde , Niv Cohen