中文
相关论文

相关论文: Auditing Games for Sandbagging

200 篇论文

In response to rising concerns surrounding the safety, security, and trustworthiness of Generative AI (GenAI) models, practitioners and regulators alike have pointed to AI red-teaming as a key component of their strategies for identifying…

计算机与社会 · 计算机科学 2024-08-29 Michael Feffer , Anusha Sinha , Wesley Hanwen Deng , Zachary C. Lipton , Hoda Heidari

As jailbreaks, adversarially crafted inputs that bypass safety constraints, continue to be discovered in Large Language Models, practitioners increasingly rely on fine-tuning as a defensive strategy. Yet the theoretical foundations…

机器学习 · 计算机科学 2026-05-20 Paul Wang , Jade Garcia-Bourrée , Anne-Marie Kermarrec , Vincent Corruble

Despite considerable efforts on making them robust, real-world AI-based systems remain vulnerable to decision based attacks, as definitive proofs of their operational robustness have so far proven intractable. Canonical robustness…

Compiler diagnostics for type inference failures are notoriously bad, and type classes only make the problem worse. By introducing a complex search process during inference, type classes can lead to wholly inscrutable or useless errors. We…

编程语言 · 计算机科学 2025-04-29 Gavin Gray , Will Crichton , Shriram Krishnamurthi

Deep Learning has already been successfully applied to analyze industrial sensor data in a variety of relevant use cases. However, the opaque nature of many well-performing methods poses a major obstacle for real-world deployment.…

机器学习 · 计算机科学 2023-10-20 Thomas Decker , Michael Lebacher , Volker Tresp

The evaluation of constitutive models, especially for high-risk and high-regret engineering applications, requires efficient and rigorous third-party calibration, validation and falsification. While there are numerous efforts to develop…

信号处理 · 电气工程与系统科学 2020-12-02 Kun Wang , WaiChing Sun , Qiang Du

Collaborative AI experimentation in industry and academia requires environments that support rapid trials while maintaining controlled access, organisational isolation, and traceable workflows. Although interest in AI sandboxes is…

A major bottleneck in characterizing the failure modes of generative AI systems is the cost and time of annotation and evaluation. Consequently, adaptive testing paradigms have gained popularity, where one opportunistically decides which…

人工智能 · 计算机科学 2026-05-11 Siyu Zhou , Patrick Vossler , Venkatesh Sivaraman , Yifan Mai , Jean Feng

Explanations for AI models in high-stakes domains like medicine often lack verifiability, which can hinder trust. To address this, we propose an interactive agent that produces explanations through an auditable sequence of actions. The…

人工智能 · 计算机科学 2025-11-04 Yuhang Huang , Zekai Lin , Fan Zhong , Lei Liu

Interpretable AI tools are often motivated by the goal of understanding model behavior in out-of-distribution (OOD) contexts. Despite the attention this area of study receives, there are comparatively few cases where these tools have…

机器学习 · 计算机科学 2023-09-25 Stephen Casper , Yuxiao Li , Jiawei Li , Tong Bu , Kevin Zhang , Kaivalya Hariharan , Dylan Hadfield-Menell

Artificial Intelligence (AI) Auditability is a core requirement for achieving responsible AI system design. However, it is not yet a prominent design feature in current applications. Existing AI auditing tools typically lack integration…

计算机与社会 · 计算机科学 2024-06-21 Laura Waltersdorfer , Fajar J. Ekaputra , Tomasz Miksa , Marta Sabou

Data-trained predictive models see widespread use, but for the most part they are used as black boxes which output a prediction or score. It is therefore hard to acquire a deeper understanding of model behavior, and in particular how…

Explainable Artificial Intelligence (XAI) is a young but very promising field of research. Unfortunately, the progress in this field is currently slowed down by divergent and incompatible goals. We separate various threads tangled within…

人工智能 · 计算机科学 2024-07-30 Przemyslaw Biecek , Wojciech Samek

We introduce WildTeaming, an automatic LLM safety red-teaming framework that mines in-the-wild user-chatbot interactions to discover 5.7K unique clusters of novel jailbreak tactics, and then composes multiple tactics for systematic…

Red teaming has evolved from its origins in military applications to become a widely adopted methodology in cybersecurity and AI. In this paper, we take a critical look at the practice of AI red teaming. We argue that despite its current…

人工智能 · 计算机科学 2025-11-03 Subhabrata Majumdar , Brian Pendleton , Abhishek Gupta

As machine learning algorithms increasingly influence critical decision making in different application areas, understanding human strategic behavior in response to these systems becomes vital. We explore individuals' choice between…

机器学习 · 计算机科学 2026-03-17 Sura Alhanouti , Parinaz Naghizadeh

Embodied AI has made significant progress acting in unexplored environments. However, tasks such as object search have largely focused on efficient policy learning. In this work, we identify several gaps in current search methods: They…

机器人学 · 计算机科学 2025-01-15 Sai Prasanna , Daniel Honerkamp , Kshitij Sirohi , Tim Welschehold , Wolfram Burgard , Abhinav Valada

We tackle a fundamental problem in empirical game-theoretic analysis (EGTA), that of learning equilibria of simulation-based games. Such games cannot be described in analytical form; instead, a black-box simulator can be queried to obtain…

计算机科学与博弈论 · 计算机科学 2019-06-03 Enrique Areyan Viqueira , Cyrus Cousins , Eli Upfal , Amy Greenwald

As part of an effort to apply the rigorous guarantees of formal verification to multi-agent systems, the field of equilibrium analysis, also called rational verification, studies equilibria in multiplayer games to reason about system-level…

计算机科学与博弈论 · 计算机科学 2026-04-28 Senthil Rajasekaran , Jean-François Raskin , Moshe Y. Vardi

To determine the safety of large language models (LLMs), AI developers must be able to assess their dangerous capabilities. But simple prompting strategies often fail to elicit an LLM's full capabilities. One way to elicit capabilities more…

机器学习 · 计算机科学 2024-05-31 Ryan Greenblatt , Fabien Roger , Dmitrii Krasheninnikov , David Krueger