中文
相关论文

相关论文: ShieldGemma 2: Robust and Tractable Image Content …

200 篇论文

Social media platforms are being increasingly used by malicious actors to share unsafe content, such as images depicting sexual activity, cyberbullying, and self-harm. Consequently, major platforms use artificial intelligence (AI) and human…

计算机视觉与模式识别 · 计算机科学 2024-01-23 Mazal Bethany , Brandon Wherry , Nishant Vishwamitra , Peyman Najafirad

The rapid development of image generation models has facilitated the widespread dissemination of generated images on social networks, creating favorable conditions for provably secure image steganography. However, existing methods face…

多媒体 · 计算机科学 2024-12-18 Yuang Qi , Kejiang Chen , Na Zhao , Zijin Yang , Weiming Zhang

The ideal AI safety moderation system would be both structurally interpretable (so its decisions can be reliably explained) and steerable (to align to safety standards and reflect a community's values), which current systems fall short on.…

As Large Language Models (LLMs) are increasingly deployed in safety-critical applications, robust content moderation becomes essential. We present a comprehensive evaluation of 14 open-source safety guard models on a curated benchmark of…

计算与语言 · 计算机科学 2026-05-29 Reetu Raj Harsh , Bhaskarjit Sarmah , Stefano Pasquali

Large Language Models (LLMs) have demonstrated exceptional performance across various tasks, but their security vulnerabilities can be exploited by attackers to generate harmful content, causing adverse impacts across various societal…

密码学与安全 · 计算机科学 2025-12-17 Fan Yang

Large Language Model (LLM) safety guardrail models have emerged as a primary defense mechanism against harmful content generation, yet their robustness against sophisticated adversarial attacks remains poorly characterized. This study…

密码学与安全 · 计算机科学 2025-12-01 Richard J. Young

Currently, large models are prone to generating harmful content when faced with complex attack instructions, significantly reducing their defensive capabilities. To address this issue, this paper proposes a method based on constructing data…

密码学与安全 · 计算机科学 2025-01-03 Keke Zhai

Vision Language Models (VLMs) are increasingly deployed in autonomous vehicles and mobile systems, making it crucial to evaluate their ability to support safer decision-making in complex environments. However, existing benchmarks…

计算机视觉与模式识别 · 计算机科学 2026-01-14 Takara Taniguchi , Kuniaki Saito , Atsushi Hashimoto

Jailbreak attacks on multimodal AI systems remain underexplored, even though unsafe image generation can have more severe consequences than unsafe text and current defenses are relatively immature. We introduce PAST2HARM, a simple yet…

计算与语言 · 计算机科学 2026-05-28 Snehasis Mukhopadhyay

The jailbreak attack can bypass the safety measures of a Large Language Model (LLM), generating harmful content. This misuse of LLM has led to negative societal consequences. Currently, there are two main approaches to address jailbreak…

计算与语言 · 计算机科学 2024-03-25 Zezhong Wang , Fangkai Yang , Lu Wang , Pu Zhao , Hongru Wang , Liang Chen , Qingwei Lin , Kam-Fai Wong

Recent advances in text-to-image diffusion models have enabled the creation of a new form of digital art: optical illusions--visual tricks that create different perceptions of reality. However, adversaries may misuse such techniques to…

密码学与安全 · 计算机科学 2025-07-31 Yiting Qu , Ziqing Yang , Yihan Ma , Michael Backes , Savvas Zannettou , Yang Zhang

Social media platforms struggle to protect users from harmful content through content moderation. These platforms have recently leveraged machine learning models to cope with the vast amount of user-generated content daily. Since moderation…

机器学习 · 计算机科学 2023-01-27 Donghyun Son , Byounggyu Lew , Kwanghee Choi , Yongsu Baek , Seungwoo Choi , Beomjun Shin , Sungjoo Ha , Buru Chang

Large language models (LLMs) can still be jailbroken into producing harmful outputs despite safety alignment. Existing attacks show this vulnerability, but not the internal mechanisms that cause it. This study asks whether jailbreak success…

计算与语言 · 计算机科学 2026-04-28 Nilanjana Das , Manas Gaur

Adversarial attacks, particularly the Fast Gradient Sign Method (FGSM) and Projected Gradient Descent (PGD) pose significant threats to the robustness of deep learning models in image classification. This paper explores and refines defense…

密码学与安全 · 计算机科学 2025-05-15 Hetvi Waghela , Jaydip Sen , Sneha Rakshit

Vision-Language Models (VLMs) face significant safety vulnerabilities from malicious prompt attacks due to weakened alignment during visual integration. Existing defenses suffer from efficiency and robustness. To address these challenges,…

机器学习 · 计算机科学 2026-04-09 Peigui Qi , Kunsheng Tang , Yanpu Yu , Jialin Wu , Yide Song , Wenbo Zhou , Zhicong Huang , Cheng Hong , Weiming Zhang , Nenghai Yu

Diffusion-based text-to-image models have shown immense potential for various image-related tasks. However, despite their prominence and popularity, customizing these models using unauthorized data also brings serious privacy and…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Sen Peng , Jijia Yang , Mingyue Wang , Jianfei He , Xiaohua Jia

Background: Cyber-attacks have evolved rapidly in recent years, many individuals and business owners have been affected by cyber-attacks in various ways. Cyber-attacks include various threats such as ransomware, malware, phishing, and…

密码学与安全 · 计算机科学 2026-01-13 Keerthi Kumar. M , Swarun Kumar Joginpelly , Sunil Khemka , Lakshmi. S R , Navin Chhibber

Large language models are increasingly used for mental health support, yet their conversational coherence alone does not ensure clinical appropriateness. Existing general-purpose safeguards often fail to distinguish between therapeutic…

Vision Large Language Models (VLLMs) represent a significant advancement in artificial intelligence by integrating image-processing capabilities with textual understanding, thereby enhancing user interactions and expanding application…

计算与语言 · 计算机科学 2025-05-09 Madhur Jindal , Saurabh Deshpande

Advanced text-to-image models such as DALL$\cdot$E 2 and Midjourney possess the capacity to generate highly realistic images, raising significant concerns regarding the potential proliferation of unsafe content. This includes adult,…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Zhongjie Ba , Jieming Zhong , Jiachen Lei , Peng Cheng , Qinglong Wang , Zhan Qin , Zhibo Wang , Kui Ren