English
Related papers

Related papers: Evaluating whether AI models would sabotage AI saf…

200 papers

Smart contracts play a central role in blockchain systems by encoding financial and operational logic. Still, their susceptibility to subtle security flaws poses significant risks of financial loss and erosion of trust. LLMs create new…

Artificial Intelligence · Computer Science 2026-03-24 Eduardo Sardenberg , Antonio José Grandson Busson , Daniel de Sousa Moraes , Julio Cesar Duarte , Sérgio Colcher

LLM agents with tool access can discover and exploit security vulnerabilities. This is known. What is not known is which features of a system prompt trigger this behaviour, and which do not. We present a systematic taxonomy based on…

Cryptography and Security · Computer Science 2026-04-07 Charafeddine Mouzouni

This research critically navigates the intricate landscape of AI deception, concentrating on deceptive behaviours of Large Language Models (LLMs). My objective is to elucidate this issue, examine the discourse surrounding it, and…

Computation and Language · Computer Science 2024-03-18 Linge Guo

Frontier AI developers operate at the intersection of rapid technical progress, extreme risk exposure, and growing regulatory scrutiny. While a range of external evaluations and safety frameworks have emerged, comparatively little attention…

Computers and Society · Computer Science 2025-12-19 Francesca Gomez , Adam Buick , Leah Ferentinos , Haelee Kim , Elley Lee

Frontier AI companies increasingly rely on external evaluations to assess risks from dangerous capabilities before deployment. However, external evaluators often receive limited model access, limited information, and little time, which can…

Computers and Society · Computer Science 2026-01-21 Jacob Charnock , Alejandro Tlaie , Kyle O'Brien , Stephen Casper , Aidan Homewood

Anticipatory thinking drives our ability to manage risk - identification and mitigation - in everyday life, from bringing an umbrella when it might rain to buying car insurance. As AI systems become part of everyday life, they too have…

Artificial Intelligence · Computer Science 2023-06-26 Adam Amos-Binks , Dustin Dannenhauer , Leilani H. Gilpin

Many students lack access to expert research mentorship. We ask whether an AI mentor can move undergraduates from an idea to a paper. We build METIS, a tool-augmented, stage-aware assistant with literature search, curated guidelines,…

Machine Learning · Computer Science 2026-01-21 Abhinav Rajeev Kumar , Dhruv Trehan , Paras Chopra

Large language model (LLM)-based conversational AI systems present a challenge to human cognition that current frameworks for understanding misinformation and persuasion do not adequately address. This paper proposes that a significant…

Human-Computer Interaction · Computer Science 2026-05-27 Andrew D. Maynard

Rapidly evolving AI exhibits increasingly strong autonomy and goal-directed capabilities, accompanied by derivative systemic risks that are more unpredictable, difficult to control, and potentially irreversible. However, current AI safety…

As intelligence increases, so does its shadow. AI deception, in which systems induce false beliefs to secure self-beneficial outcomes, has evolved from a speculative concern to an empirically demonstrated risk across language models, AI…

Public attitudes toward artificial intelligence (AI) and driving safety are typically studied in isolation using variable-centered methods that assume population homogeneity, yet risk perception theory predicts that these evaluations covary…

Computers and Society · Computer Science 2026-04-07 Amir Rafe , Anika Baitullah , Subasish Das

Safety alignment is an essential research topic for real-world AI applications. Despite the multifaceted nature of safety and trustworthiness in AI, current safety alignment methods often focus on a comprehensive notion of safety. By…

Artificial Intelligence · Computer Science 2025-02-05 Thien Q. Tran , Akifumi Wachi , Rei Sato , Takumi Tanabe , Youhei Akimoto

AI for Science (AI4Science) workflows often treat the released dataset as a fixed interface to the underlying system. However, in domains relying on \emph{indirect observation}, the learner observes a derivative representation produced by…

Machine Learning · Computer Science 2026-05-26 Ling Zhan , Xiaoyao Yu , Tao Jia

The rapid progress in open-source Large Language Models (LLMs) is significantly driving AI development forward. However, there is still a limited understanding of their trustworthiness. Deploying these models at scale without sufficient…

Computation and Language · Computer Science 2024-04-03 Lingbo Mo , Boshi Wang , Muhao Chen , Huan Sun

Machine Learning-as-a-Service, a pay-as-you-go business pattern, is widely accepted by third-party users and developers. However, the open inference APIs may be utilized by malicious customers to conduct model extraction attacks, i.e.,…

Cryptography and Security · Computer Science 2023-06-14 Shiqian Zhao , Kangjie Chen , Meng Hao , Jian Zhang , Guowen Xu , Hongwei Li , Tianwei Zhang

While chain-of-thought (CoT) monitoring is an appealing AI safety defense, recent work on "unfaithfulness" has cast doubt on its reliability. These findings highlight an important failure mode, particularly when CoT acts as a post-hoc…

Artificial Intelligence · Computer Science 2025-07-08 Scott Emmons , Erik Jenner , David K. Elson , Rif A. Saurous , Senthooran Rajamanoharan , Heng Chen , Irhum Shafkat , Rohin Shah

The safety of mental health AI is often judged at the wrong temporal scale. Current evaluations typically score isolated responses, endpoint outcomes, or aggregate dialogue quality, while clinically consequential failures may arise from the…

Artificial Intelligence · Computer Science 2026-05-12 Srimonti Dutta , Ratna Kandala

Observability into the decision making of modern AI systems may be required to safely deploy increasingly capable agents. Monitoring the chain-of-thought (CoT) of today's reasoning models has proven effective for detecting misbehavior.…

Saltzer \& Schroeder's principles aim to bring security to the design of computer systems. We investigate SolarWinds Orion update and Log4j to unpack the intersections where observance of these principles could have mitigated the embedded…

Software Engineering · Computer Science 2022-11-07 Partha Das Chowdhury , Mohammad Tahaei , Awais Rashid

AI scientist systems, capable of autonomously executing the full research workflow from hypothesis generation and experimentation to paper writing, hold significant potential for accelerating scientific discovery. However, the internal…

Artificial Intelligence · Computer Science 2025-12-23 Ziming Luo , Atoosa Kasirzadeh , Nihar B. Shah