English
Related papers

Related papers: Auditing Games for Sandbagging

200 papers

In response to rising concerns surrounding the safety, security, and trustworthiness of Generative AI (GenAI) models, practitioners and regulators alike have pointed to AI red-teaming as a key component of their strategies for identifying…

Computers and Society · Computer Science 2024-08-29 Michael Feffer , Anusha Sinha , Wesley Hanwen Deng , Zachary C. Lipton , Hoda Heidari

As jailbreaks, adversarially crafted inputs that bypass safety constraints, continue to be discovered in Large Language Models, practitioners increasingly rely on fine-tuning as a defensive strategy. Yet the theoretical foundations…

Machine Learning · Computer Science 2026-05-20 Paul Wang , Jade Garcia-Bourrée , Anne-Marie Kermarrec , Vincent Corruble

Despite considerable efforts on making them robust, real-world AI-based systems remain vulnerable to decision based attacks, as definitive proofs of their operational robustness have so far proven intractable. Canonical robustness…

Artificial Intelligence · Computer Science 2025-05-07 Ilias Tsingenopoulos , Vera Rimmer , Davy Preuveneers , Fabio Pierazzi , Lorenzo Cavallaro , Wouter Joosen

Compiler diagnostics for type inference failures are notoriously bad, and type classes only make the problem worse. By introducing a complex search process during inference, type classes can lead to wholly inscrutable or useless errors. We…

Programming Languages · Computer Science 2025-04-29 Gavin Gray , Will Crichton , Shriram Krishnamurthi

Deep Learning has already been successfully applied to analyze industrial sensor data in a variety of relevant use cases. However, the opaque nature of many well-performing methods poses a major obstacle for real-world deployment.…

Machine Learning · Computer Science 2023-10-20 Thomas Decker , Michael Lebacher , Volker Tresp

The evaluation of constitutive models, especially for high-risk and high-regret engineering applications, requires efficient and rigorous third-party calibration, validation and falsification. While there are numerous efforts to develop…

Signal Processing · Electrical Eng. & Systems 2020-12-02 Kun Wang , WaiChing Sun , Qiang Du

Collaborative AI experimentation in industry and academia requires environments that support rapid trials while maintaining controlled access, organisational isolation, and traceable workflows. Although interest in AI sandboxes is…

A major bottleneck in characterizing the failure modes of generative AI systems is the cost and time of annotation and evaluation. Consequently, adaptive testing paradigms have gained popularity, where one opportunistically decides which…

Artificial Intelligence · Computer Science 2026-05-11 Siyu Zhou , Patrick Vossler , Venkatesh Sivaraman , Yifan Mai , Jean Feng

Explanations for AI models in high-stakes domains like medicine often lack verifiability, which can hinder trust. To address this, we propose an interactive agent that produces explanations through an auditable sequence of actions. The…

Artificial Intelligence · Computer Science 2025-11-04 Yuhang Huang , Zekai Lin , Fan Zhong , Lei Liu

Interpretable AI tools are often motivated by the goal of understanding model behavior in out-of-distribution (OOD) contexts. Despite the attention this area of study receives, there are comparatively few cases where these tools have…

Machine Learning · Computer Science 2023-09-25 Stephen Casper , Yuxiao Li , Jiawei Li , Tong Bu , Kevin Zhang , Kaivalya Hariharan , Dylan Hadfield-Menell

Artificial Intelligence (AI) Auditability is a core requirement for achieving responsible AI system design. However, it is not yet a prominent design feature in current applications. Existing AI auditing tools typically lack integration…

Computers and Society · Computer Science 2024-06-21 Laura Waltersdorfer , Fajar J. Ekaputra , Tomasz Miksa , Marta Sabou

Data-trained predictive models see widespread use, but for the most part they are used as black boxes which output a prediction or score. It is therefore hard to acquire a deeper understanding of model behavior, and in particular how…

Explainable Artificial Intelligence (XAI) is a young but very promising field of research. Unfortunately, the progress in this field is currently slowed down by divergent and incompatible goals. We separate various threads tangled within…

Artificial Intelligence · Computer Science 2024-07-30 Przemyslaw Biecek , Wojciech Samek

We introduce WildTeaming, an automatic LLM safety red-teaming framework that mines in-the-wild user-chatbot interactions to discover 5.7K unique clusters of novel jailbreak tactics, and then composes multiple tactics for systematic…

Red teaming has evolved from its origins in military applications to become a widely adopted methodology in cybersecurity and AI. In this paper, we take a critical look at the practice of AI red teaming. We argue that despite its current…

Artificial Intelligence · Computer Science 2025-11-03 Subhabrata Majumdar , Brian Pendleton , Abhishek Gupta

As machine learning algorithms increasingly influence critical decision making in different application areas, understanding human strategic behavior in response to these systems becomes vital. We explore individuals' choice between…

Machine Learning · Computer Science 2026-03-17 Sura Alhanouti , Parinaz Naghizadeh

Embodied AI has made significant progress acting in unexplored environments. However, tasks such as object search have largely focused on efficient policy learning. In this work, we identify several gaps in current search methods: They…

We tackle a fundamental problem in empirical game-theoretic analysis (EGTA), that of learning equilibria of simulation-based games. Such games cannot be described in analytical form; instead, a black-box simulator can be queried to obtain…

Computer Science and Game Theory · Computer Science 2019-06-03 Enrique Areyan Viqueira , Cyrus Cousins , Eli Upfal , Amy Greenwald

As part of an effort to apply the rigorous guarantees of formal verification to multi-agent systems, the field of equilibrium analysis, also called rational verification, studies equilibria in multiplayer games to reason about system-level…

Computer Science and Game Theory · Computer Science 2026-04-28 Senthil Rajasekaran , Jean-François Raskin , Moshe Y. Vardi

To determine the safety of large language models (LLMs), AI developers must be able to assess their dangerous capabilities. But simple prompting strategies often fail to elicit an LLM's full capabilities. One way to elicit capabilities more…

Machine Learning · Computer Science 2024-05-31 Ryan Greenblatt , Fabien Roger , Dmitrii Krasheninnikov , David Krueger
‹ Prev 1 4 5 6 7 8 10 Next ›