English
Related papers

Related papers: Scheming in the wild: detecting real-world AI sche…

200 papers

Intent detection is a crucial component of modern conversational systems, since accurately identifying user intent at the beginning of a conversation is essential for generating effective responses. Recent efforts have focused on studying…

Computation and Language · Computer Science 2025-09-09 Liang Zhang , Yuan Li , Shijie Zhang , Zheng Zhang , Xitong Li

Explainable Artificial Intelligence (XAI) enhances the transparency and interpretability of AI models, addressing their inherent opacity. In cybersecurity, particularly within the Internet of Medical Things (IoMT), the black-box nature of…

Cryptography and Security · Computer Science 2025-09-16 Mohammed Yacoubi , Omar Moussaoui , C. Drocourt

When AI agents operating with access to sensitive information encounter a conflict between completing an assigned task and following rules or ethical constraints, they can resort to unsanctioned behaviour. Existing inference time safety…

Cryptography and Security · Computer Science 2026-05-01 Francesca Gomez

Alignment research on large language models (LLMs) increasingly depends on understanding how these systems are used in everyday contexts. Yet naturalistic interaction data is difficult to access due to privacy constraints and platform…

Human-Computer Interaction · Computer Science 2026-03-24 Cathy Mengying Fang , Sheer Karny , Chayapatr Archiwaranguprok , Yasith Samaradivakara , Pat Pataranutaporn , Pattie Maes

Modern AI agents execute real-world side effects through tool calls such as file operations, shell commands, HTTP requests, and database queries. A single unsafe action, including accidental deletion, credential exposure, or data…

Artificial Intelligence · Computer Science 2026-05-07 Chenglin Yang

Out-of-scope intent detection is of practical importance in task-oriented dialogue systems. Since the distribution of outlier utterances is arbitrary and unknown in the training stage, existing methods commonly rely on strong assumptions on…

Computation and Language · Computer Science 2021-06-18 Li-Ming Zhan , Haowen Liang , Bo Liu , Lu Fan , Xiao-Ming Wu , Albert Y. S. Lam

As autonomous and agentic AI systems scale in robotic and human-machine environments, managing hallucination and persistent but unjustified action remains an open challenge. Rather than attributing these failures solely to model or…

Artificial Intelligence · Computer Science 2026-05-28 Srini Ramaswamy

As AI systems become more advanced, concerns about large-scale risks from misuse or accidents have grown. This report analyzes the technical research into safe AI development being conducted by three leading AI companies: Anthropic, Google…

Computers and Society · Computer Science 2024-09-26 Oscar Delaney , Oliver Guest , Zoe Williams

The AI Incident Database was inspired by aviation safety databases, which enable collective learning from failures to prevent future incidents. The database documents hundreds of AI failures, collected from the news and media. However,…

Computers and Society · Computer Science 2025-05-08 Isabel Richards , Claire Benn , Miri Zilka

This paper develops an economic model of the Open Source Intelligence (OSINT) attention economy in contemporary armed conflict. We conceptualize attention (e.g. social media views, followers, likes) as revenue, and time and risk spent in…

General Economics · Economics 2025-09-16 Jonathan Teagan

Large Language Models (LLMs) and generative search systems are increasingly used for information seeking by diverse populations with varying preferences for knowledge sourcing and presentation. While users can customize LLM behavior through…

Human-Computer Interaction · Computer Science 2025-12-12 Nicholas Clark , Ryan Bai , Tanu Mitra

The multilingual nature of the Internet increases complications in the cybersecurity community's ongoing efforts to strategically mine threat intelligence from OSINT data on the web. OSINT sources such as social media, blogs, and dark web…

Computation and Language · Computer Science 2018-07-20 Priyanka Ranade , Sudip Mittal , Anupam Joshi , Karuna Joshi

The Signalgate incident of March 2025, wherein senior US national security officials inadvertently disclosed sensitive military operational details via the encrypted messaging platform Signal, highlights critical vulnerabilities in…

Cryptography and Security · Computer Science 2025-10-29 Paul Benjamin Lowry , Gregory D. Moody , Robert Willison , Clay Posey

Large language model (LLM) developers aim for their models to be honest, helpful, and harmless. However, when faced with malicious requests, models are trained to refuse, sacrificing helpfulness. We show that frontier LLMs can develop a…

AI systems in high-consequence domains such as defense, intelligence, and disaster response must detect rare, high-impact events while operating under tight resource constraints. Traditional annotation strategies that prioritize label…

Machine Learning · Computer Science 2025-05-22 Dave Cook , Tim Klawa

Social media engagement prediction is a central challenge in computational social science, particularly for understanding how users interact with misinformation. Existing approaches often treat engagement as a homogeneous time-series…

Social and Information Networks · Computer Science 2026-02-03 Lin Tian , Marian-Andrei Rizoiu

The argument for persistent social media influence campaigns, often funded by malicious entities, is gaining traction. These entities utilize instrumented profiles to disseminate divisive content and disinformation, shaping public…

Computers and Society · Computer Science 2024-01-26 Hina Qayyum , Muhammad Ikram , Benjamin Zi Hao Zhao , an D. Wood , Nicolas Kourtellis , Mohamed Ali Kaafar

With artificial intelligence (AI) being applied to bring autonomy to decision-making in safety-critical domains such as the ones typified in the aerospace and emergency-response services, there has been a call to address the ethical…

Artificial Intelligence · Computer Science 2025-09-03 Julian Gerald Dcruz , Argyrios Zolotas , Niall Ross Greenwood , Miguel Arana-Catania

Chain-of-thought (CoT) monitoring is one of the most promising tools we have for detecting model misbehavior, but its effectiveness depends on models faithfully externalizing their reasoning. Motivated by this vulnerability, we study…

Machine Learning · Computer Science 2026-05-18 Reilly Haskins , Bilal Chughtai , Joshua Engels

There has been a proliferation of media reports about so-called AI psychosis in the last year. Not surprisingly, this has prompted growing academic work on the ways in which AI chatbots such as ChatGPT, Claude, and Replika might aggravate…

Human-Computer Interaction · Computer Science 2026-05-27 Kasper Møller Nielsen , Lucy Osler