中文
相关论文

相关论文: Evaluating whether AI models would sabotage AI saf…

200 篇论文

Smart contracts play a central role in blockchain systems by encoding financial and operational logic. Still, their susceptibility to subtle security flaws poses significant risks of financial loss and erosion of trust. LLMs create new…

LLM agents with tool access can discover and exploit security vulnerabilities. This is known. What is not known is which features of a system prompt trigger this behaviour, and which do not. We present a systematic taxonomy based on…

密码学与安全 · 计算机科学 2026-04-07 Charafeddine Mouzouni

This research critically navigates the intricate landscape of AI deception, concentrating on deceptive behaviours of Large Language Models (LLMs). My objective is to elucidate this issue, examine the discourse surrounding it, and…

计算与语言 · 计算机科学 2024-03-18 Linge Guo

Frontier AI developers operate at the intersection of rapid technical progress, extreme risk exposure, and growing regulatory scrutiny. While a range of external evaluations and safety frameworks have emerged, comparatively little attention…

计算机与社会 · 计算机科学 2025-12-19 Francesca Gomez , Adam Buick , Leah Ferentinos , Haelee Kim , Elley Lee

Frontier AI companies increasingly rely on external evaluations to assess risks from dangerous capabilities before deployment. However, external evaluators often receive limited model access, limited information, and little time, which can…

计算机与社会 · 计算机科学 2026-01-21 Jacob Charnock , Alejandro Tlaie , Kyle O'Brien , Stephen Casper , Aidan Homewood

Anticipatory thinking drives our ability to manage risk - identification and mitigation - in everyday life, from bringing an umbrella when it might rain to buying car insurance. As AI systems become part of everyday life, they too have…

人工智能 · 计算机科学 2023-06-26 Adam Amos-Binks , Dustin Dannenhauer , Leilani H. Gilpin

Many students lack access to expert research mentorship. We ask whether an AI mentor can move undergraduates from an idea to a paper. We build METIS, a tool-augmented, stage-aware assistant with literature search, curated guidelines,…

机器学习 · 计算机科学 2026-01-21 Abhinav Rajeev Kumar , Dhruv Trehan , Paras Chopra

Large language model (LLM)-based conversational AI systems present a challenge to human cognition that current frameworks for understanding misinformation and persuasion do not adequately address. This paper proposes that a significant…

人机交互 · 计算机科学 2026-05-27 Andrew D. Maynard

Rapidly evolving AI exhibits increasingly strong autonomy and goal-directed capabilities, accompanied by derivative systemic risks that are more unpredictable, difficult to control, and potentially irreversible. However, current AI safety…

As intelligence increases, so does its shadow. AI deception, in which systems induce false beliefs to secure self-beneficial outcomes, has evolved from a speculative concern to an empirically demonstrated risk across language models, AI…

Public attitudes toward artificial intelligence (AI) and driving safety are typically studied in isolation using variable-centered methods that assume population homogeneity, yet risk perception theory predicts that these evaluations covary…

计算机与社会 · 计算机科学 2026-04-07 Amir Rafe , Anika Baitullah , Subasish Das

Safety alignment is an essential research topic for real-world AI applications. Despite the multifaceted nature of safety and trustworthiness in AI, current safety alignment methods often focus on a comprehensive notion of safety. By…

人工智能 · 计算机科学 2025-02-05 Thien Q. Tran , Akifumi Wachi , Rei Sato , Takumi Tanabe , Youhei Akimoto

AI for Science (AI4Science) workflows often treat the released dataset as a fixed interface to the underlying system. However, in domains relying on \emph{indirect observation}, the learner observes a derivative representation produced by…

机器学习 · 计算机科学 2026-05-26 Ling Zhan , Xiaoyao Yu , Tao Jia

The rapid progress in open-source Large Language Models (LLMs) is significantly driving AI development forward. However, there is still a limited understanding of their trustworthiness. Deploying these models at scale without sufficient…

计算与语言 · 计算机科学 2024-04-03 Lingbo Mo , Boshi Wang , Muhao Chen , Huan Sun

Machine Learning-as-a-Service, a pay-as-you-go business pattern, is widely accepted by third-party users and developers. However, the open inference APIs may be utilized by malicious customers to conduct model extraction attacks, i.e.,…

密码学与安全 · 计算机科学 2023-06-14 Shiqian Zhao , Kangjie Chen , Meng Hao , Jian Zhang , Guowen Xu , Hongwei Li , Tianwei Zhang

While chain-of-thought (CoT) monitoring is an appealing AI safety defense, recent work on "unfaithfulness" has cast doubt on its reliability. These findings highlight an important failure mode, particularly when CoT acts as a post-hoc…

The safety of mental health AI is often judged at the wrong temporal scale. Current evaluations typically score isolated responses, endpoint outcomes, or aggregate dialogue quality, while clinically consequential failures may arise from the…

人工智能 · 计算机科学 2026-05-12 Srimonti Dutta , Ratna Kandala

Observability into the decision making of modern AI systems may be required to safely deploy increasingly capable agents. Monitoring the chain-of-thought (CoT) of today's reasoning models has proven effective for detecting misbehavior.…

Saltzer \& Schroeder's principles aim to bring security to the design of computer systems. We investigate SolarWinds Orion update and Log4j to unpack the intersections where observance of these principles could have mitigated the embedded…

软件工程 · 计算机科学 2022-11-07 Partha Das Chowdhury , Mohammad Tahaei , Awais Rashid

AI scientist systems, capable of autonomously executing the full research workflow from hypothesis generation and experimentation to paper writing, hold significant potential for accelerating scientific discovery. However, the internal…

人工智能 · 计算机科学 2025-12-23 Ziming Luo , Atoosa Kasirzadeh , Nihar B. Shah