English
Related papers

Related papers: Scheming in the wild: detecting real-world AI sche…

200 papers

Out of Scope (OOS) detection in Conversational AI solutions enables a chatbot to handle a conversation gracefully when it is unable to make sense of the end-user query. Accurately tagging a query as out-of-domain is particularly hard in…

Computation and Language · Computer Science 2023-01-18 Cheng Qian , Haode Qi , Gengyu Wang , Ladislav Kunc , Saloni Potdar

As AI agents become more widely deployed, we are likely to see an increasing number of incidents: events involving AI agent use that directly or indirectly cause harm. For example, agents could be prompt-injected to exfiltrate private…

Computers and Society · Computer Science 2025-08-21 Carson Ezell , Xavier Roberts-Gaal , Alan Chan

This work investigates the potential of torrent metadata as a source for open-source intelligence (OSINT), with a focus on user profiling and behavioral analysis. While peer-to-peer (P2P) networks such as BitTorrent are well studied with…

Information Retrieval · Computer Science 2026-01-06 Annelies de Jong , Giuseppe Cascavilla , Jessica De Pascale

Prompt-based interfaces for Large Language Models (LLMs) have made prototyping and building AI-powered applications easier than ever before. However, identifying potential harms that may arise from AI applications remains a challenge,…

Human-Computer Interaction · Computer Science 2024-07-03 Zijie J. Wang , Chinmay Kulkarni , Lauren Wilcox , Michael Terry , Michael Madaio

Large Language Models and commercial speech synthesis systems now enable highly realistic AI-generated voice scams (vishing), raising urgent concerns about deception at scale. Yet it remains unclear whether individuals can reliably…

Cryptography and Security · Computer Science 2026-03-27 Zoha Hayat Bhatti , Bakhtawar Ahtisham , Seemal Tausif , Niklas George , Nida ul Habib Bajwa , Mobin Javed

We introduce HonestCyberEval, a new benchmark for assessing AI models' capabilities and risks in automated software exploitation, focusing on their ability to detect and exploit vulnerabilities in real-world software systems. Our evaluation…

Cryptography and Security · Computer Science 2025-08-27 Dan Ristea , Vasilios Mavroudis

The democratization of generative AI introduces new forms of human-AI interaction and raises urgent safety, ethical, and cybersecurity concerns. We develop a socio-technical explanation for how generative AI enables and scales cybercrime.…

Computers and Society · Computer Science 2025-12-04 Truong Jack Luu , Binny M. Samuel

The rapid growth of social media presents a unique opportunity to study coordinated agent behavior in an unfiltered environment. Online processes often exhibit complex structures that reflect the nature of the user behavior, whether it is…

Social and Information Networks · Computer Science 2025-06-17 Anna Kalenkova , Lewis Mitchell , Ethan Johnson

Data breaches have begun to take on new dimensions and their prediction is becoming of great importance to organizations. Prior work has addressed this issue mainly from a technical perspective and neglected other interfering aspects such…

Cryptography and Security · Computer Science 2024-11-20 Hicham Hammouchi , Narjisse Nejjari , Ghita Mezzour , Mounir Ghogho , Houda Benbrahim

The rapid scaling of AI has spurred a growing emphasis on ethical considerations in both development and practice. This has led to the formulation of increasingly sophisticated model auditing and reporting requirements, as well as…

Human-Computer Interaction · Computer Science 2024-10-01 Mahasweta Chakraborti , Bert Joseph Prestoza , Nicholas Vincent , Seth Frey

While existing social bot detectors perform well on benchmarks, their robustness across diverse real-world scenarios remains limited due to unclear ground truth and varied misleading cues. In particular, the impact of shortcut learning,…

Computation and Language · Computer Science 2026-03-24 Shiyan Zheng , Herun Wan , Minnan Luo , Junhang Huang

Scaling up language models has significantly increased their capabilities. But larger models are slower models, and so there is now an extensive body of work (e.g., speculative sampling or parallel decoding) that improves the (average case)…

Cryptography and Security · Computer Science 2024-10-23 Nicholas Carlini , Milad Nasr

Incident monitoring can drive safety improvements in high-reliability industries and population-scale technologies, but remains underdeveloped in AI governance. Public databases catalog thousands of AI incidents, but simple incident counts…

Computers and Society · Computer Science 2026-05-08 Isaak Mengesha , Branwen Owen , Charlie Collins , Tina Wong , Simon Mylius , Peter Slattery , Sean McGregor

Output from generative AI such as ChatGPT, can be repetitive and biased. But more worrying is that this output can mysteriously tip mid-response from good (correct) to bad (misleading or wrong) without the user noticing. In 2024 alone, this…

Artificial Intelligence · Computer Science 2025-08-05 Neil F. Johnson , Frank Yingjie Huo

Human oversight of AI is promoted as a safeguard against risks such as inaccurate outputs, system malfunctions, or violations of fundamental rights, and is mandated in regulation like the European AI Act. Yet debates on human oversight have…

Cryptography and Security · Computer Science 2026-03-06 Jonas C. Ditz , Veronika Lazar , Elmar Lichtmeß , Carola Plesch , Matthias Heck , Kevin Baum , Markus Langer

As autonomous driving technology matures, safety and robustness of its key components, including trajectory prediction, is vital. Though real-world datasets, such as Waymo Open Motion, provide realistic recorded scenarios for model…

Robotics · Computer Science 2024-02-06 Benjamin Stoler , Ingrid Navarro , Meghdeep Jana , Soonmin Hwang , Jonathan Francis , Jean Oh

As autonomous agents become more capable of performing real-world tasks, distinguishing scheming behavior from benign task pursuit may become a central AI control problem. Existing monitors often rely on chain-of-thought access or internal…

Computation and Language · Computer Science 2026-05-29 Aditya Sinha , Akshat Naik , Victor Gillioz , Simon Storf , Kilian Merkelbach , Rich Barton-Cooper , Axel Højmark , Marius Hobbhahn

AI scientist systems, capable of autonomously executing the full research workflow from hypothesis generation and experimentation to paper writing, hold significant potential for accelerating scientific discovery. However, the internal…

Artificial Intelligence · Computer Science 2025-12-23 Ziming Luo , Atoosa Kasirzadeh , Nihar B. Shah

AI systems have become increasingly capable of dangerous behaviours in many domains. This raises the question: Do models sometimes choose to violate human instructions in order to perform behaviour that is more useful for certain goals? We…

Artificial Intelligence · Computer Science 2026-05-08 Jonas Wiedermann-Möller , Leonard Dung , Maksym Andriushchenko

Email phishing remains a prevalent cyber threat, targeting victims to extract sensitive information or deploy malicious software. This paper explores the integration of open-source intelligence (OSINT) tools and machine learning (ML) models…

Cryptography and Security · Computer Science 2025-01-16 Panharith An , Rana Shafi , Tionge Mughogho , Onyango Allan Onyango