English
Related papers

Related papers: Red-Teaming for Generative AI: Silver Bullet or Se…

200 papers

Generative AI, in particular text-based "foundation models" (large models trained on a huge variety of information including the internet), can generate speech that could be problematic under a wide range of liability regimes. Machine…

Computers and Society · Computer Science 2023-08-21 Peter Henderson , Tatsunori Hashimoto , Mark Lemley

AI safety is a rapidly growing area of research that seeks to prevent the harm and misuse of frontier AI technology, particularly with respect to generative AI (GenAI) tools that are capable of creating realistic and high-quality content…

Artificial Intelligence · Computer Science 2025-02-19 Pin-Yu Chen

Generative Artificial Intelligence (GenAI) enables digital representatives to make decisions on behalf of team members in collaborative tasks, but faces challenges in accurately representing preferences. While supplying GenAI with detailed…

Computer Science and Game Theory · Computer Science 2025-12-16 Kexin Chen , Jianwei Huang , Yuan Luo

Large Language Model (LLM) safeguards, which implement request refusals, have become a widely adopted mitigation strategy against misuse. At the intersection of adversarial machine learning and AI safety, safeguard red teaming has…

Cryptography and Security · Computer Science 2025-06-10 Zifan Wang , Christina Q. Knight , Jeremy Kritz , Willow E. Primack , Julian Michael

What makes safety claims about general purpose AI systems such as large language models trustworthy? We show that rather than the capabilities of security tools such as alignment and red teaming procedures, it is security practices based on…

Cryptography and Security · Computer Science 2025-07-30 Petr Spelda , Vit Stritecky

Since 2022, generative AI (genAI) has rapidly become integrated into workplaces. Though organizations have made commitments to use this technology "responsibly", how organizations and their employees prioritize responsibility in their…

Computers and Society · Computer Science 2025-01-29 Genevieve Smith , Natalia Luka , Merrick Osborne , Brian Lattimore , Jessica Newman , Brandie Nonnecke , Brent Mittelstadt

Red-teaming, where adversarial prompts are crafted to expose harmful behaviors and assess risks, offers a dynamic approach to surfacing underlying stereotypical bias in large language models. Because such subtle harms are best recognized by…

Human-Computer Interaction · Computer Science 2026-02-24 Sieun Kim , Yeeun Jo , Sungmin Na , Hyunseung Lim , Eunchae Lee , Yu Min Choi , Soohyun Cho , Hwajung Hong

Generative Artificial Intelligence (GenAI) is becoming ubiquitous in our daily lives. The increase in computational power and data availability has led to a proliferation of both single- and multi-modal models. As the GenAI ecosystem…

Generative AI systems produce a range of risks. To ensure the safety of generative AI systems, these risks must be evaluated. In this paper, we make two main contributions toward establishing such evaluations. First, we propose a…

Generative AI models are capable of performing a wide variety of tasks that have traditionally required creativity and human understanding. During training, they learn patterns from existing data and can subsequently generate new content…

The rise of Generative AI (GenAI) brings about transformative potential across sectors, but its dual-use nature also amplifies risks. Governments globally are grappling with the challenge of regulating GenAI, balancing innovation against…

Large-scale pre-trained generative models are taking the world by storm, due to their abilities in generating creative content. Meanwhile, safeguards for these generative models are developed, to protect users' rights and safety, most of…

Cryptography and Security · Computer Science 2024-10-14 Guanlin Li , Kangjie Chen , Shudong Zhang , Jie Zhang , Tianwei Zhang

The rapid growth of Large Language Models (LLMs) presents significant privacy, security, and ethical concerns. While much research has proposed methods for defending LLM systems against misuse by malicious actors, researchers have recently…

Computation and Language · Computer Science 2025-03-06 Alberto Purpura , Sahil Wadhwa , Jesse Zymet , Akshay Gupta , Andy Luo , Melissa Kazemi Rad , Swapnil Shinde , Mohammad Shahed Sorower

As generative AI (GenAI) agents become more common in enterprise settings, they introduce security challenges that differ significantly from those posed by traditional systems. These agents are not just LLMs; they reason, remember, and act,…

Cryptography and Security · Computer Science 2025-05-06 Vineeth Sai Narajala , Om Narayan

Game theory offers a powerful framework for analyzing strategic interactions among decision-makers, providing tools to model, analyze, and predict their behavior. However, implementing game theory can be challenging due to difficulties in…

Computer Science and Game Theory · Computer Science 2024-05-15 Yaoqi Yang , Hongyang Du , Geng Sun , Zehui Xiong , Dusit Niyato , Zhu Han

Recent developments in AI governance and safety research have called for red-teaming methods that can effectively surface potential risks posed by AI models. Many of these calls have emphasized how the identities and backgrounds of…

Artificial Intelligence · Computer Science 2025-10-28 Wesley Hanwen Deng , Sunnie S. Y. Kim , Akshita Jha , Ken Holstein , Motahhare Eslami , Lauren Wilcox , Leon A Gatys

To evaluate the safety and usefulness of deployment protocols for untrusted AIs, AI Control uses a red-teaming exercise played between a protocol designer and an adversary. This paper introduces AI-Control Games, a formal decision-making…

Artificial Intelligence · Computer Science 2026-05-08 Charlie Griffin , Louis Thomson , Buck Shlegeris , Alessandro Abate

This whitepaper highlights the dual importance of securing generative AI (genAI) platforms and leveraging genAI for cybersecurity. As genAI technologies proliferate, their misuse poses significant risks, including data breaches, model…

Cryptography and Security · Computer Science 2024-10-21 Hari Hayagreevan , Souvik Khamaru

Generative Artificial Intelligence (GenAI) presents significant advancements but also introduces novel security challenges, particularly within agentic workflows where AI agents operate autonomously. These risks escalate in multi-agent…

Cryptography and Security · Computer Science 2025-06-24 Sunil Kumar Jang Bahadur , Gopala Dhar

Red teaming is critical for identifying vulnerabilities and building trust in current LLMs. However, current automated methods for Large Language Models (LLMs) rely on brittle prompt templates or single-turn attacks, failing to capture the…

Machine Learning · Computer Science 2025-08-07 Roman Belaire , Arunesh Sinha , Pradeep Varakantham