English
Related papers

Related papers: Evaluating whether AI models would sabotage AI saf…

200 papers

This technical report presents methods developed by the UK AI Security Institute for assessing whether advanced AI systems reliably follow intended goals. Specifically, we evaluate whether frontier models sabotage safety research when…

Artificial Intelligence · Computer Science 2026-04-02 Alexandra Souly , Robert Kirk , Jacob Merizian , Abby D'Cruz , Xander Davies

Sufficiently capable models could subvert human oversight and decision-making in important contexts. For example, in the context of AI development, models could covertly sabotage efforts to evaluate their own dangerous capabilities, to…

Recent work has demonstrated the plausibility of frontier AI models scheming -- knowingly and covertly pursuing an objective misaligned with its developer's intentions. Such behavior could be very hard to detect, and if present in future…

Machine Learning · Computer Science 2025-07-04 Mary Phuong , Roland S. Zimmermann , Ziyue Wang , David Lindner , Victoria Krakovna , Sarah Cogan , Allan Dafoe , Lewis Ho , Rohin Shah

AI systems are increasingly able to autonomously conduct realistic software engineering tasks, and may soon be deployed to automate machine learning (ML) R&D itself. Frontier AI systems may be deployed in safety-critical settings, including…

Frontier models are increasingly trained and deployed as autonomous agent. One safety concern is that AI agents might covertly pursue misaligned goals, hiding their true capabilities and objectives - also known as scheming. We study whether…

Artificial Intelligence · Computer Science 2025-01-16 Alexander Meinke , Bronson Schoen , Jérémy Scheurer , Mikita Balesni , Rusheb Shah , Marius Hobbhahn

As AI systems are increasingly used to conduct research autonomously, misaligned systems could introduce subtle flaws that produce misleading results while evading detection. We introduce Auditing Sabotage Bench, a benchmark for evaluating…

Artificial Intelligence · Computer Science 2026-04-28 Eric Gan , Aryan Bhatt , Buck Shlegeris , Julian Stastny , Vivek Hebbar

AI coding scaffolds like Claude Code and Codex use retrying: blocking actions flagged as risky and continuing the trajectory. We study retrying from an AI control perspective, which treats the model as potentially adversarial. We find that…

Artificial Intelligence · Computer Science 2026-05-27 James Lucassen , Adam Kaufman

AI leaders and safety reports increasingly warn that advances in model reasoning may enable biological misuse, including by low-expertise users, while major labs describe safeguards as expanding but still evolving rather than settled. This…

Computers and Society · Computer Science 2026-04-24 Michael Richter

As Large Language Models (LLMs) are increasingly deployed as autonomous agents in complex and long horizon settings, it is critical to evaluate their ability to sabotage users by pursuing hidden objectives. We study the ability of frontier…

Trustworthy evaluations of dangerous capabilities are increasingly crucial for determining whether an AI system is safe to deploy. One empirically demonstrated threat is sandbagging - the strategic underperformance on evaluations by AI…

Cryptography and Security · Computer Science 2025-11-03 Chloe Li , Mary Phuong , Noah Y. Siegel

As foundation models grow increasingly more intelligent, reliable and trustworthy safety evaluation becomes more indispensable than ever. However, an important question arises: Whether and how an advanced AI system would perceive the…

Artificial Intelligence · Computer Science 2026-03-16 Yihe Fan , Wenqi Zhang , Xudong Pan , Min Yang

We sketch how developers of frontier AI systems could construct a structured rationale -- a 'safety case' -- that an AI system is unlikely to cause catastrophic outcomes through scheming. Scheming is a potential threat model where AI…

Trustworthy capability evaluations are crucial for ensuring the safety of AI systems, and are becoming a key component of AI regulation. However, the developers of an AI system, or the AI system itself, may have incentives for evaluations…

Artificial Intelligence · Computer Science 2025-02-10 Teun van der Weij , Felix Hofstätter , Ollie Jaffe , Samuel F. Brown , Francis Rhys Ward

Large language models are trained to refuse harmful requests, but can they accurately predict when they will refuse before responding? We investigate this question through a systematic study where models first predict their refusal…

Computation and Language · Computer Science 2026-04-02 Tanay Gondil

The rapid advancement of AI systems has raised widespread concerns about potential harms of frontier AI systems and the need for responsible evaluation and oversight. In this position paper, we argue that frontier AI companies should report…

Computers and Society · Computer Science 2025-03-25 Dillon Bowen , Ann-Kathrin Dombrowski , Adam Gleave , Chris Cundy

We present swarm-attack, an open-source adversarial testing framework in which multiple lightweight LLM agents coordinate through shared memory, parallel exploration, and evolutionary optimization. Together, our results demonstrate that…

Cryptography and Security · Computer Science 2026-05-12 Michael A. Riegler , Inga Strümke

Today's leading AI models engage in sophisticated behaviour when placed in strategic competition. They spontaneously attempt deception, signaling intentions they do not intend to follow; they demonstrate rich theory of mind, reasoning about…

Artificial Intelligence · Computer Science 2026-02-17 Kenneth Payne

Frontier artificial intelligence (AI) systems pose increasing risks to society, making it essential for developers to provide assurances about their safety. One approach to offering such assurances is through a safety case: a structured,…

Computers and Society · Computer Science 2024-11-14 Arthur Goemans , Marie Davidsen Buhl , Jonas Schuett , Tomek Korbak , Jessica Wang , Benjamin Hilton , Geoffrey Irving

Mitigating reward hacking--where AI systems misbehave due to flaws or misspecifications in their learning objectives--remains a key challenge in constructing capable and aligned models. We show that we can monitor a frontier reasoning…

Artificial Intelligence · Computer Science 2025-03-18 Bowen Baker , Joost Huizinga , Leo Gao , Zehao Dou , Melody Y. Guan , Aleksander Madry , Wojciech Zaremba , Jakub Pachocki , David Farhi

Artificial intelligence (AI) control protocols assume that trusted large language model (LLM) monitors reliably assess proposed actions across all deployment contexts. This paper tests that assumption in the geographic dimension. We audit…

Computers and Society · Computer Science 2026-04-16 Jason Hung
‹ Prev 1 2 3 10 Next ›