中文
相关论文

相关论文: Red Team Redemption: A Structured Comparison of Op…

200 篇论文

Adversarial examples are perturbed inputs designed to fool machine learning models. Adversarial training injects such examples into training data to increase robustness. To scale this technique to large datasets, perturbations are crafted…

Existing efforts in safeguarding LLMs are limited in actively exposing the vulnerabilities of the target LLM and readily adapting to newly emerging safety risks. To address this, we present Purple-teaming LLMs with Adversarial Defender…

计算与语言 · 计算机科学 2024-07-03 Jingyan Zhou , Kun Li , Junan Li , Jiawen Kang , Minda Hu , Xixin Wu , Helen Meng

To address the problem of imperfect confrontation strategy caused by the lack of information of game environment in the simulation of non-complete information dynamic countermeasure modeling for intelligent game, the hierarchical analysis…

人工智能 · 计算机科学 2022-03-30 Xiangri Lu , Hongbin Ma , Zhanqing Wang

Gemini is increasingly used to perform tasks on behalf of users, where function-calling and tool-use capabilities enable the model to access user data. Some tools, however, require access to untrusted data introducing risk. Adversaries can…

Security operators use red teams to simulate real attackers and proactively find defense gaps. In realistic enterprise settings, this involves executing multi-host network attacks spanning many "stepping stone" hosts. Unfortunately, red…

密码学与安全 · 计算机科学 2025-11-25 Brian Singer , Keane Lucas , Lakshmi Adiga , Meghna Jain , Lujo Bauer , Vyas Sekar

We introduce a three stage pipeline: resized-diverse-inputs (RDIM), diversity-ensemble (DEM) and region fitting, that work together to generate transferable adversarial examples. We first explore the internal relationship between existing…

计算机视觉与模式识别 · 计算机科学 2021-12-14 Junhua Zou , Zhisong Pan , Junyang Qiu , Xin Liu , Ting Rui , Wei Li

Adversarial robustness has emerged as an important topic in deep learning as carefully crafted attack samples can significantly disturb the performance of a model. Many recent methods have proposed to improve adversarial robustness by…

机器学习 · 计算机科学 2019-08-08 Hao-Yun Chen , Jhao-Hong Liang , Shih-Chieh Chang , Jia-Yu Pan , Yu-Ting Chen , Wei Wei , Da-Cheng Juan

This paper provides an efficient computational scheme to handle general security games from an adversarial risk analysis perspective. Two cases in relation to single-stage and multi-stage simultaneous defend-attack games motivate our…

计算机科学与博弈论 · 计算机科学 2025-06-04 Jose Manuel Camacho , Roi Naveiro , David Rios Insua

Recent studies have revealed the vulnerability of pre-trained language models to adversarial attacks. Existing adversarial defense techniques attempt to reconstruct adversarial examples within feature or text spaces. However, these methods…

计算与语言 · 计算机科学 2024-04-02 Heng Yang , Ke Li

Our objective in this paper is to develop a machinery that makes a given organizational strategic plan resilient to the actions of competitor agents (adverse environmental actions). We assume that we are given a goal tree representing…

人工智能 · 计算机科学 2025-02-18 Aditya Ghose , Asjad Khan

Artificial intelligence (AI) is being ubiquitously adopted to automate processes in science and industry. However, due to its often intricate and opaque nature, AI has been shown to possess inherent vulnerabilities which can be maliciously…

密码学与安全 · 计算机科学 2023-12-20 Mathew J. Walter , Aaron Barrett , Kimberly Tam

Prompt injection poses serious security risks to real-world LLM applications, particularly autonomous agents. Although many defenses have been proposed, their robustness against adaptive attacks remains insufficiently evaluated, potentially…

机器学习 · 计算机科学 2026-03-16 Chenlong Yin , Runpeng Geng , Yanting Wang , Jinyuan Jia

As generative AI technologies find more and more real-world applications, the importance of testing their performance and safety seems paramount. "Red-teaming" has quickly become the primary approach to test AI models--prioritized by AI…

计算机与社会 · 计算机科学 2026-01-09 Tarleton Gillespie , Ryland Shaw , Mary L. Gray , Jina Suh

Despite extensive diagnostics and debugging by developers, AI systems sometimes exhibit harmful unintended behaviors. Finding and fixing these is challenging because the attack surface is so large -- it is not tractable to exhaustively…

密码学与安全 · 计算机科学 2025-07-30 Stephen Casper , Lennart Schulze , Oam Patel , Dylan Hadfield-Menell

With the increasing growth of cyber-attack incidences, it is important to develop innovative and effective techniques to assess and defend networked systems against cyber attacks. One of the well-known techniques for this is performing…

密码学与安全 · 计算机科学 2021-05-19 Simon Yusuf Enoch , Zhibin Huang , Chun Yong Moon , Donghwan Lee , Myung Kil Ahn , Dong Seong Kim

The rapid integration of Generative AI (GenAI) into various applications necessitates robust risk management strategies which includes Red Teaming (RT) - an evaluation method for simulating adversarial attacks. Unfortunately, RT for GenAI…

密码学与安全 · 计算机科学 2025-05-01 Tam n. Nguyen

AI coding assistants like GitHub Copilot are rapidly transforming software development, but their safety remains deeply uncertain-especially in high-stakes domains like cybersecurity. Current red-teaming tools often rely on fixed benchmarks…

Adversarial attacks have emerged as a major challenge to the trustworthy deployment of machine learning models, particularly in computer vision applications. These attacks have a varied level of potency and can be implemented in both white…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Nandish Chattopadhyay , Abdul Basit , Bassem Ouni , Muhammad Shafique

As generative AI, particularly large language models (LLMs), become increasingly integrated into production applications, new attack surfaces and vulnerabilities emerge and put a focus on adversarial threats in natural language and…

Cybersecurity operations demand assistant LLMs that support diverse workflows without exposing sensitive data. Existing solutions either rely on proprietary APIs with privacy risks or on open models lacking domain adaptation. To bridge this…

密码学与安全 · 计算机科学 2026-03-10 Naufal Suryanto , Muzammal Naseer , Pengfei Li , Syed Talal Wasim , Jinhui Yi , Juergen Gall , Paolo Ceravolo , Ernesto Damiani