English
Related papers

Related papers: Red Teaming AI Red Teaming

200 papers

Existing evaluations of AI misuse safeguards provide a patchwork of evidence that is often difficult to connect to real-world decisions. To bridge this gap, we describe an end-to-end argument (a "safety case") that misuse safeguards reduce…

Machine Learning · Computer Science 2025-05-26 Joshua Clymer , Jonah Weinbaum , Robert Kirk , Kimberly Mai , Selena Zhang , Xander Davies

Artificial intelligence (AI) is increasingly being used to augment and automate cyber operations, altering the scale, speed, and accessibility of malicious activity. These shifts raise urgent questions about when AI systems introduce…

Cryptography and Security · Computer Science 2026-01-27 Krystal Jackson , Deepika Raman , Jessica Newman , Nada Madkour , Charlotte Yuan , Evan R. Murphy

With the proliferation of AI-enabled software systems in smart manufacturing, the role of such systems moves away from a reactive to a proactive role that provides context-specific support to manufacturing operators. In the frame of the EU…

Software Engineering · Computer Science 2022-01-24 Philipp Haindl , Georg Buchgeher , Maqbool Khan , Bernhard Moser

Scientific research organizations that are developing and deploying Artificial Intelligence (AI) systems are at the intersection of technological progress and ethical considerations. The push for Responsible AI (RAI) in such institutions…

Artificial Intelligence · Computer Science 2023-12-18 Muneera Bano , Didar Zowghi , Pip Shea , Georgina Ibarra

The increasing use of AI technologies has led to increasing AI incidents, posing risks and causing harm to individuals, organizations, and society. This study recognizes and addresses the lack of standardized protocols for reliably and…

Computers and Society · Computer Science 2025-01-28 Avinash Agarwal , Manisha J Nene

In the rapidly evolving field of cybersecurity, ensuring the reproducibility of AI-driven research is critical to maintaining the reliability and integrity of security systems. This paper addresses the reproducibility crisis within the…

Machine Learning · Computer Science 2024-12-17 Richard H. Moulton , Gary A. McCully , John D. Hastings

Despite the intricacies involved in designing a computer as a teampartner, we can observe patterns in team behavior which allow us to describe at a general level how AI systems are to collaborate with humans. Whereas most work on…

Human-Computer Interaction · Computer Science 2021-01-18 Jurriaan van Diggelen , Wiard Jorritsma , Bob van der Vecht

Artificial Intelligence (AI) has the opportunity to revolutionize the way the United States Department of Defense (DoD) and Intelligence Community (IC) address the challenges of evolving threats, data deluge, and rapid courses of action.…

Artificial Intelligence · Computer Science 2019-05-10 Vijay Gadepally , Justin Goodwin , Jeremy Kepner , Albert Reuther , Hayley Reynolds , Siddharth Samsi , Jonathan Su , David Martinez

AI practitioners typically strive to develop the most accurate systems, making an implicit assumption that the AI system will function autonomously. However, in practice, AI systems often are used to provide advice to people in domains…

Artificial Intelligence · Computer Science 2021-02-23 Gagan Bansal , Besmira Nushi , Ece Kamar , Eric Horvitz , Daniel S. Weld

Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language processing tasks, but their vulnerability to jailbreak attacks poses significant security risks. This survey paper presents a comprehensive analysis…

Computation and Language · Computer Science 2024-12-18 Tarun Raheja , Nilay Pochhi , F. D. C. M. Curie

A key task in AI practice is to assess potential impacts to prevent harm. Current AI tools assisting AI impact assessment have not been designed or evaluated for collaborative team brainstorming, and they do not capture the range of views…

Human-Computer Interaction · Computer Science 2026-05-01 Jarod Govers , Sanja Šćepanović , Daniele Quercia

As artificial intelligence (AI) becomes deeply embedded in critical services and everyday products, it is increasingly exposed to security threats which traditional cyber defenses were not designed to handle. In this paper, we investigate…

Cryptography and Security · Computer Science 2026-03-06 Natalia Krawczyk , Mateusz Szczepkowski , Adrian Brodzik , Krzysztof Bocianiak

Adversarial testing of large language models (LLMs) is crucial for their safe and responsible deployment. We introduce a novel approach for automated generation of adversarial evaluation datasets to test the safety of LLM generations on new…

Software Engineering · Computer Science 2023-12-01 Bhaktipriya Radharapu , Kevin Robinson , Lora Aroyo , Preethi Lahoti

Machine learning and software development share processes and methodologies for reliably delivering products to customers. This work proposes the use of a new teaming construct for forming machine learning teams for better combatting…

Machine Learning · Computer Science 2021-10-22 Josh Kalin , David Noever , Matthew Ciolino

Embedding artificial intelligence into systems introduces significant challenges to modern engineering practices. Hazard analysis tools and processes have not yet been adequately adapted to the new paradigm. This paper describes initial…

Software Engineering · Computer Science 2022-03-30 Nikolas Martelaro , Carol J. Smith , Tamara Zilovic

The dual offensive and defensive utility of Large Language Models (LLMs) highlights a critical gap in AI security: the lack of unified frameworks for dynamic, iterative adversarial adaptation hardening. To bridge this gap, we propose the…

Cryptography and Security · Computer Science 2026-01-28 Lige Huang , Zicheng Liu , Jie Zhang , Lewen Yan , Dongrui Liu , Jing Shao

Language-conditioned robot models have the potential to enable robots to perform a wide range of tasks based on natural language instructions. However, assessing their safety and effectiveness remains challenging because it is difficult to…

As large language models (LLMs) become increasingly prevalent across many real-world applications, understanding and enhancing their robustness to adversarial attacks is of paramount importance. Existing methods for identifying adversarial…

Warning: this paper contains content that may be inappropriate or offensive. As generative models become available for public use in various applications, testing and analyzing vulnerabilities of these models has become a priority. In this…

Artificial Intelligence · Computer Science 2024-11-11 Ninareh Mehrabi , Palash Goyal , Christophe Dupuy , Qian Hu , Shalini Ghosh , Richard Zemel , Kai-Wei Chang , Aram Galstyan , Rahul Gupta

Collaborative human-AI (HAI) teaming combines the unique skills and capabilities of humans and machines in sustained teaming interactions leveraging the strengths of each. In tasks involving regular exposure to novelty and uncertainty,…

Human-Computer Interaction · Computer Science 2024-04-03 Melanie J. McGrath , Andreas Duenser , Justine Lacey , Cecile Paris