English
Related papers

Related papers: Scheming in the wild: detecting real-world AI sche…

200 papers

Online Social Networks (OSNs), such as Facebook, provide users with tools to share information along with a set of privacy controls preferences to regulate the spread of information. Current privacy controls are efficient to protect content…

Social and Information Networks · Computer Science 2016-11-22 Giuseppe Cascavilla , Filipe Beato , Andrea Burattin , Mauro Conti , Luigi Vincenzo Mancini

Recent work has demonstrated the plausibility of frontier AI models scheming -- knowingly and covertly pursuing an objective misaligned with its developer's intentions. Such behavior could be very hard to detect, and if present in future…

Machine Learning · Computer Science 2025-07-04 Mary Phuong , Roland S. Zimmermann , Ziyue Wang , David Lindner , Victoria Krakovna , Sarah Cogan , Allan Dafoe , Lewis Ho , Rohin Shah

We sketch how developers of frontier AI systems could construct a structured rationale -- a 'safety case' -- that an AI system is unlikely to cause catastrophic outcomes through scheming. Scheming is a potential threat model where AI…

Highly capable AI systems could secretly pursue misaligned goals -- what we call "scheming". Because a scheming AI would deliberately try to hide its misaligned goals and actions, measuring and mitigating scheming requires different…

Frontier models are increasingly trained and deployed as autonomous agent. One safety concern is that AI agents might covertly pursue misaligned goals, hiding their true capabilities and objectives - also known as scheming. We study whether…

Artificial Intelligence · Computer Science 2025-01-16 Alexander Meinke , Bronson Schoen , Jérémy Scheurer , Mikita Balesni , Rusheb Shah , Marius Hobbhahn

Small businesses need vulnerability assessments to identify and mitigate cyber risks. Cybersecurity clinics provide a solution by offering students hands-on experience while delivering free vulnerability assessments to local organizations.…

Human-Computer Interaction · Computer Science 2025-02-21 Anirban Mukhopadhyay , Kurt Luther

Open Source Intelligence (OSINT) refers to intelligence efforts based on freely available data. It has become a frequent topic of conversation on social media, where private users or networks can share their findings. Such data is highly…

Social and Information Networks · Computer Science 2024-09-04 Johannes Niu , Mila Stillman , Philipp Seeberger , Anna Kruspe

Facebook represents the current de-facto choice for social media, changing the nature of social relationships. The increasing amount of personal information that runs through this platform publicly exposes user behaviour and social trends,…

Cryptography and Security · Computer Science 2019-11-01 Aimilia Panagiotou , Bogdan Ghita , Stavros Shiaeles , Keltoum Bendiab

Open Source Intelligence (OSINT) investigations, which rely entirely on publicly available data such as social media, play an increasingly important role in solving crimes and holding governments accountable. The growing volume of data and…

Human-Computer Interaction · Computer Science 2025-02-21 Anirban Mukhopadhyay , Sukrit Venkatagiri , Kurt Luther

Intent Detection systems in the real world are exposed to complexities of imbalanced datasets containing varying perception of intent, unintended correlations and domain-specific aberrations. To facilitate benchmarking which can reflect…

Computation and Language · Computer Science 2021-03-25 Gaurav Arora , Chirag Jain , Manas Chaturvedi , Krupal Modi

This paper examines the role of Open Source Intelligence (OSINT) on Twitter regarding the Russo-Ukrainian war, distinguishing between genuine OSINT and deceptive misinformation efforts, termed "BULLSHINT." Utilizing a dataset spanning from…

Social and Information Networks · Computer Science 2025-08-06 Johannes Niu , Mila Stillman , Anna Kruspe

As large language models (LLMs) are increasingly deployed as interactive agents, open-ended human-AI interactions can involve deceptive behaviors with serious real-world consequences, yet existing evaluations remain largely…

Artificial Intelligence · Computer Science 2026-02-09 Yichen Wu , Qianqian Gao , Xudong Pan , Geng Hong , Min Yang

As frontier language models are increasingly deployed as autonomous agents pursuing complex, long-term objectives, there is increased risk of scheming: agents covertly pursuing misaligned goals. Prior work has focused on showing agents are…

Artificial Intelligence · Computer Science 2026-03-31 Mia Hopman , Jannes Elstner , Maria Avramidou , Amritanshu Prasad , David Lindner

Artificial intelligence systems are now deployed at scale across sectors, accompanied by a growing number of real-world incidents ranging from misinformation and cybercrime to autonomous-system failures. Databases of AI incidents index…

Computers and Society · Computer Science 2026-04-23 Sophia Abraham , Taiye Chen , Cyril Chhun , Giovanna Jaramillo-Gutierrez , Simon Mylius , Sayash Raaj , Peter Slattery , Sean McGregor

High-risk industries like nuclear and aviation use real-time monitoring to detect dangerous system conditions. Similarly, Large Language Models (LLMs) need monitoring safeguards. We propose a real-time framework to predict harmful AI…

Artificial Intelligence · Computer Science 2025-05-21 Maheep Chaudhary , Fazl Barez

As artificial intelligence (AI) systems become increasingly deployed across the world, they are also increasingly implicated in AI incidents - harm events to individuals and society. As a result, industry, civil society, and governments…

Computers and Society · Computer Science 2024-09-26 Kevin Paeth , Daniel Atherton , Nikiforos Pittaras , Heather Frase , Sean McGregor

This report examines whether advanced AIs that perform well in training will be doing so in order to gain power later -- a behavior I call "scheming" (also sometimes called "deceptive alignment"). I conclude that scheming is a disturbingly…

Computers and Society · Computer Science 2023-11-29 Joe Carlsmith

We introduce scheming honeypot evaluations, a framework for testing whether models will pursue instrumental goals if given the opportunity. Our scheming honeypot evaluations take the form of coding tasks in Google's alignment research…

Machine Learning · Computer Science 2026-05-29 Victoria Krakovna , David Lindner , Lewis Ho , Sebastian Farquhar , Rohin Shah

Two years after publicly launching the AI Incident Database (AIID) as a collection of harms or near harms produced by AI in the world, a backlog of "issues" that do not meet its incident ingestion criteria have accumulated in its review…

Computers and Society · Computer Science 2022-11-21 Sean McGregor , Kevin Paeth , Khoa Lam

We systematically evaluate the quality of widely used adversarial safety datasets from two perspectives: in isolation and in practice. In isolation, we examine how well these datasets reflect real-world adversarial attacks based on three…

Cryptography and Security · Computer Science 2026-04-24 Shahriar Golchin , Marc Wetter
‹ Prev 1 2 3 10 Next ›