English
Related papers

Related papers: Games for AI Control: Models of Safety Evaluations…

200 papers

As AI systems become more capable and widely deployed as agents, ensuring their safe operation becomes critical. AI control offers one approach to mitigating the risk from untrusted AI agents by monitoring their actions and intervening or…

Artificial Intelligence · Computer Science 2025-11-06 Jon Kutasov , Chloe Loughridge , Yuqi Sun , Henry Sleight , Buck Shlegeris , Tyler Tracy , Joe Benton

Technology development efforts in autonomy and cyber-defense have been evolving independently of each other, over the past decade. In this paper, we report our ongoing effort to integrate these two presently distinct areas into a single…

Computer Science and Game Theory · Computer Science 2020-02-07 Mohamadreza Ahmadi , Arun A. Viswanathan , Michel D. Ingham , Kymie Tan , Aaron D. Ames

Selecting the combination of security controls that will most effectively protect a system's assets is a difficult task. If the wrong controls are selected, the system may be left vulnerable to cyber-attacks that can impact the…

Cryptography and Security · Computer Science 2024-10-31 Dylan Léveillé , Jason Jaskolka

The AI Control research agenda aims to develop control protocols: safety techniques that prevent untrusted AI systems from taking harmful actions during deployment. Because human oversight is expensive, one approach is trusted monitoring,…

Cryptography and Security · Computer Science 2026-02-12 Ashwin Sreevatsa , Sebastian Prasanna , Cody Rushing

This paper studies a stochastic game theoretic approach to security and intrusion detection in communication and computer networks. Specifically, an Attacker and a Defender take part in a two-player game over a network of nodes whose…

Cryptography and Security · Computer Science 2010-03-15 Kien C. Nguyen , Tansu Alpcan , Tamer Basar

Red teaming has evolved from its origins in military applications to become a widely adopted methodology in cybersecurity and AI. In this paper, we take a critical look at the practice of AI red teaming. We argue that despite its current…

Artificial Intelligence · Computer Science 2025-11-03 Subhabrata Majumdar , Brian Pendleton , Abhishek Gupta

Network systems often contain vulnerabilities that remain unfixed in a network for various reasons, such as the lack of a patch or knowledge to fix them. With the presence of such residual vulnerabilities, the network administrator should…

Cryptography and Security · Computer Science 2022-11-04 Narges Khakpour , David Parker

In response to rising concerns surrounding the safety, security, and trustworthiness of Generative AI (GenAI) models, practitioners and regulators alike have pointed to AI red-teaming as a key component of their strategies for identifying…

Computers and Society · Computer Science 2024-08-29 Michael Feffer , Anusha Sinha , Wesley Hanwen Deng , Zachary C. Lipton , Hoda Heidari

As a transformative general-purpose technology, AI has empowered various industries and will continue to shape our lives through ubiquitous applications. Despite the enormous benefits from wide-spread AI deployment, it is crucial to address…

Computer Science and Game Theory · Computer Science 2023-05-25 Na Zhang , Kun Yue , Chao Fang

Cybersecurity threats are becoming increasingly sophisticated, making traditional defense mechanisms and manual red teaming approaches insufficient for modern organizations. While red teaming has long been recognized as an effective method…

Cryptography and Security · Computer Science 2026-02-26 Shruti Srivastava , Kiranmayee Janardhan , Shaurya Jauhari

An AI control protocol is a plan for usefully deploying AI systems that aims to prevent an AI from intentionally causing some unacceptable outcome. This paper investigates how well AI systems can generate and act on their own strategies for…

Machine Learning · Computer Science 2025-04-07 Alex Mallen , Charlie Griffin , Misha Wagner , Alessandro Abate , Buck Shlegeris

What makes safety claims about general purpose AI systems such as large language models trustworthy? We show that rather than the capabilities of security tools such as alignment and red teaming procedures, it is security practices based on…

Cryptography and Security · Computer Science 2025-07-30 Petr Spelda , Vit Stritecky

Several recent works have studied the societal effects of AI; these include issues such as fairness, robustness, and safety. In many of these objectives, a learner seeks to minimize its worst-case loss over a set of predefined distributions…

Artificial Intelligence · Computer Science 2023-10-31 Yash Gupta , Runtian Zhai , Arun Suggala , Pradeep Ravikumar

This position paper argues for two claims regarding AI testing and evaluation. First, to remain informative about deployment behaviour, evaluations need account for the possibility that AI systems understand their circumstances and reason…

Computer Science and Game Theory · Computer Science 2025-08-22 Vojtech Kovarik , Eric Olav Chen , Sami Petersen , Alexis Ghersengorin , Vincent Conitzer

Evaluating the safety of AI Systems is a pressing concern for organizations deploying them. In addition to the societal damage done by the lack of fairness of those systems, deployers are concerned about the legal repercussions and the…

Stochastic games are a convenient formalism for modelling systems that comprise rational agents competing or collaborating within uncertain environments. Probabilistic model checking techniques for this class of models allow us to formally…

Logic in Computer Science · Computer Science 2022-11-14 Marta Kwiatkowska , Gethin Norman , David Parker , Gabriel Santos

Control evaluations measure whether monitoring and security protocols for AI systems prevent intentionally subversive AI models from causing harm. Our work presents the first control evaluation performed in an agent environment. We…

Machine Learning · Computer Science 2025-04-15 Aryan Bhatt , Cody Rushing , Adam Kaufman , Tyler Tracy , Vasil Georgiev , David Matolcsi , Akbir Khan , Buck Shlegeris

Red teaming has emerged as a critical practice in assessing the possible risks of AI models and systems. It aids in the discovery of novel risks, stress testing possible gaps in existing mitigations, enriching existing quantitative safety…

Computers and Society · Computer Science 2025-03-24 Lama Ahmad , Sandhini Agarwal , Michael Lampe , Pamela Mishkin

In this paper, we propose a construction scheme for a Safe-visor architecture for sandboxing unverified controllers, e.g., artificial intelligence-based (a.k.a. AI-based) controllers, in two-players non-cooperative stochastic games.…

Systems and Control · Electrical Eng. & Systems 2022-03-29 Bingzhuo Zhong , Hongpeng Cao , Majid Zamani , Marco Caccamo

Fairness is a desirable and crucial property of many protocols that handle, for instance, exchanges of message. It states that if at least one agent engaging in the protocol is honest, then either the protocol will unfold correctly and…

Computer Science and Game Theory · Computer Science 2024-11-01 Léonard Brice , Jean-François Raskin , Mathieu Sassolas , Guillaume Scerri , Marie van den Bogaard
‹ Prev 1 2 3 10 Next ›