中文
相关论文

相关论文: Password-Activated Shutdown Protocols for Misalign…

200 篇论文

The POST-Agents Proposal (PAP) is an idea for ensuring that advanced artificial agents never resist shutdown. A key part of the PAP is using a novel `Discounted Reward for Same-Length Trajectories (DReST)' reward function to train agents to…

人工智能 · 计算机科学 2026-05-13 Elliott Thornley , Alexander Roman , Christos Ziakas , Leyton Ho , Louis Thomson

Creating systems that are aligned with our goals is seen as a leading approach to create safe and beneficial AI in both leading AI companies and the academic field of AI safety. We defend the view that misaligned AGI - future, generally…

计算机与社会 · 计算机科学 2025-06-05 Max Hellrigel-Holderbaum , Leonard Dung

For over a decade, cybersecurity has relied on human labor scarcity to limit attackers to high-value targets manually or generic automated attacks at scale. Building sophisticated exploits requires deep expertise and manual effort, leading…

密码学与安全 · 计算机科学 2026-02-04 Terry Yue Zhuo , Yangruibo Ding , Wenbo Guo , Ruijie Meng

As artificial intelligence scales, the concepts of alignment, agency, and autonomy have become central to AI safety, governance, and control. However, even in human contexts, these terms lack universal definitions, varying across…

计算机与社会 · 计算机科学 2025-03-11 Krti Tallam

Autonomous AI agents powered by Large Language Models can reason, plan, and execute complex tasks, but their ability to autonomously retrieve information and run code introduces significant security risks. Existing approaches attempt to…

密码学与安全 · 计算机科学 2026-04-09 Hongyi Lu , Nian Liu , Shuai Wang , Fengwei Zhang

Reconfigurable intelligent surfaces (RISs) are arrays of passive elements that can control the reflection of the incident electromagnetic waves. While RIS are particularly useful to avoid blockages, the protocol aspects for their…

Frontier artificial intelligence (AI) systems could pose increasing risks to public safety and security. But what level of risk is acceptable? One increasingly popular approach is to define capability thresholds, which describe AI…

计算机与社会 · 计算机科学 2024-06-24 Leonie Koessler , Jonas Schuett , Markus Anderljung

Identifying the vulnerabilities of large language models (LLMs) is crucial for improving their safety by addressing inherent weaknesses. Jailbreaks, in which adversaries bypass safeguards with crafted input prompts, play a central role in…

人工智能 · 计算机科学 2026-04-03 Hamin Koo , Minseon Kim , Jaehyung Kim

The development of safety-critical systems requires the control of hazards that can potentially cause harm. To this end, safety engineers rely during the development phase on architectural solutions, called safety patterns, such as safety…

系统与控制 · 电气工程与系统科学 2020-09-23 Yuri Gil Dantas , Antoaneta Kondeva , Vivek Nigam

False data injection attacks pose a significant threat to autonomous multi-agent systems (MASs). Existing attack-resilient control strategies generally have strict assumptions on the attack signals and overlook safety constraints, such as…

系统与控制 · 电气工程与系统科学 2025-05-06 Yichao Wang , Mohamadamin Rajabinezhad , Dimitra Panagou , Shan Zuo

AI support of collaborative interactions entails mediating potential misalignment between interlocutor beliefs. Common preference alignment methods like DPO excel in static settings, but struggle in dynamic collaborative tasks where the…

计算与语言 · 计算机科学 2025-05-27 Abhijnan Nath , Carine Graff , Andrei Bachinin , Nikhil Krishnaswamy

Backdoor attacks on language models pose a significant threat to AI safety, where models behave normally on most inputs but exhibit harmful behavior when triggered by specific patterns. Detecting such backdoors through mechanistic…

计算与语言 · 计算机科学 2026-05-11 Sachin Kumar

Prior work on trustworthy AI emphasizes model-internal properties such as bias mitigation, adversarial robustness, and interpretability. As AI systems evolve into autonomous agents deployed in open environments and increasingly connected to…

人工智能 · 计算机科学 2026-05-06 Wenyue Hua , Tianyi Peng , Chi Wang , Jiaxin Pei , Ian Kaufman , Bryan Lim , Chandler Fang

Chinese authorities are extending the country's four-phase emergency response framework (prevent, warn, respond, and recover) to address risks from advanced artificial intelligence (AI). Concrete mechanisms for the proactive prevention and…

计算机与社会 · 计算机科学 2025-11-11 James Zhang , Miles Kodama , Zongze Wu , Michael Chen , Yue Zhu , Geng Hong

With software systems permeating our lives, we are entitled to expect that such systems are secure by design, and that such security endures throughout the use of these systems and their subsequent evolution. Although adaptive security…

密码学与安全 · 计算机科学 2023-06-08 Liliana Pasquale , Kushal Ramkumar , Wanling Cai , John McCarthy , Gavin Doherty , Bashar Nuseibeh

The A2AS framework is introduced as a security layer for AI agents and LLM-powered applications, similar to how HTTPS secures HTTP. A2AS enforces certified behavior, activates model self-defense, and ensures context window integrity. It…

When humans perform everyday tasks, we naturally adjust our actions based on the current state of the environment. For instance, if we intend to put something into a drawer but notice it is closed, we open it first. However, many autonomous…

机器人学 · 计算机科学 2025-08-18 Che Rin Yu , Daewon Chae , Dabin Seo , Sangwon Lee , Hyeongwoo Im , Jinkyu Kim

Autonomous coding agents are increasingly deployed as AI teammates in modern software engineering, independently authoring pull requests (PRs) that modify production code at scale. This study aims to systematically characterize how…

密码学与安全 · 计算机科学 2026-01-05 Mohammed Latif Siddiq , Xinye Zhao , Vinicius Carvalho Lopes , Beatrice Casey , Joanna C. S. Santos

Accurate local state measurement is important to ensure the reliable operation of distributed multi-agent systems (MAS). Existing fault-tolerant control strategies generally assume the sensor faults to be bounded and uncorrelated. In this…

系统与控制 · 电气工程与系统科学 2024-01-31 Shan Zuo , Yi Zhang , Yichao Wang

Jailbreak prompts pose a significant threat in AI and cybersecurity, as they are crafted to bypass ethical safeguards in large language models, potentially enabling misuse by cybercriminals. This paper analyzes jailbreak prompts from a…