中文
相关论文

相关论文: Evaluating Control Protocols for Untrusted AI Agen…

200 篇论文

What makes safety claims about general purpose AI systems such as large language models trustworthy? We show that rather than the capabilities of security tools such as alignment and red teaming procedures, it is security practices based on…

密码学与安全 · 计算机科学 2025-07-30 Petr Spelda , Vit Stritecky

The critical need for transparent and trustworthy machine learning in cybersecurity operations drives the development of this integrated Explainable AI (XAI) framework. Our methodology addresses three fundamental challenges in deploying AI…

密码学与安全 · 计算机科学 2026-02-24 Norrakith Srisumrith , Sunantha Sodsee

As artificial intelligence (AI) systems become increasingly embedded in critical societal functions, the need for robust red teaming methodologies continues to grow. In this forum piece, we examine emerging approaches to automating AI red…

计算机与社会 · 计算机科学 2025-03-31 Alice Qian Zhang , Jina Suh , Mary L. Gray , Hong Shen

AI agents deployed on decentralized infrastructures are beginning to exhibit properties that extend beyond autonomy toward what we describe as agentic sovereignty-the capacity of an operational agent to persist, act, and control resources…

计算机与社会 · 计算机科学 2026-02-17 Botao Amber Hu , Helena Rong

Artificial intelligence (AI) is being ubiquitously adopted to automate processes in science and industry. However, due to its often intricate and opaque nature, AI has been shown to possess inherent vulnerabilities which can be maliciously…

密码学与安全 · 计算机科学 2023-12-20 Mathew J. Walter , Aaron Barrett , Kimberly Tam

The rise of AI has transformed the software and hardware landscape, enabling powerful capabilities through specialized infrastructures, large-scale data storage, and advanced hardware. However, these innovations introduce unique attack…

密码学与安全 · 计算机科学 2025-08-29 Michael R Smith , Joe Ingram

The Artificial intelligence in critical sectors-healthcare, finance, and public safety-has made system integrity paramount for maintaining societal trust. Current verification methods for AI systems lack comprehensive lifecycle assurance,…

密码学与安全 · 计算机科学 2024-11-04 Mahesh Vaijainthymala Krishnamoorthy

For over a decade, cybersecurity has relied on human labor scarcity to limit attackers to high-value targets manually or generic automated attacks at scale. Building sophisticated exploits requires deep expertise and manual effort, leading…

密码学与安全 · 计算机科学 2026-02-04 Terry Yue Zhuo , Yangruibo Ding , Wenbo Guo , Ruijie Meng

AI Agents can perform complex operations at great speed, but just like all the humans we have ever hired, their intelligence remains fallible. Miscommunications aren't noticed, systemic biases have no counter-action, and inner monologues…

多智能体系统 · 计算机科学 2026-01-22 Gopal Vijayaraghavan , Prasanth Jayachandran , Arun Murthy , Sunil Govindan , Vivek Subramanian

This memorandum presents four recommendations aimed at strengthening the principles of AI model reliability and AI model governability, as DoW, ODNI, NIST, and CAISI refine AI assurance frameworks under the AI Action Plan. Our focus…

计算机与社会 · 计算机科学 2025-10-13 Matteo Pistillo , Charlotte Stix

Artificial Intelligence (AI) systems are now an integral part of multiple industries. In clinical research, AI supports automated adverse event detection in clinical trials, patient eligibility screening for protocol enrollment, and data…

人工智能 · 计算机科学 2025-12-10 Laxmiraju Kandikatla , Branislav Radeljic

Independent evaluation and red teaming are critical for identifying the risks posed by generative AI systems. However, the terms of service and enforcement strategies used by prominent AI companies to deter model misuse have disincentives…

Artificial Intelligence (AI) agents can now orchestrate cyberattacks. This development is already increasing the speed and scale of cyber attacks, decreasing attack costs, and improving the operational autonomy of cyber capabilities. To…

计算机与社会 · 计算机科学 2026-05-22 Matt Mittelsteadt , Jam Kraprayoon , Robin Staes-Polet , Oskar Galeev , Jan Wehner , Christopher Covino , Shaun Ee

A major concern amongst AI safety practitioners is the possibility of loss of control, whereby humans lose the ability to exert control over increasingly advanced AI systems. The range of concerns is wide, spanning current day risks to…

计算机与社会 · 计算机科学 2026-02-04 Steve Barrett , Anna Bruvere , Sean P. Fillingham , Catherine Rhodes , Stefano Vergani

The rapid advancement of artificial intelligence (AI) technologies presents profound challenges to societal safety. As AI systems become more capable, accessible, and integrated into critical services, the dual nature of their potential is…

人工智能 · 计算机科学 2024-12-06 Giulio Corsi , Kyle Kilian , Richard Mallah

Evolving AI systems increasingly deploy multi-agent architectures where autonomous agents collaborate, share information, and delegate tasks through developing protocols. This connectivity, while powerful, introduces novel security risks.…

密码学与安全 · 计算机科学 2025-07-30 Gauri Sharma , Vidhi Kulkarni , Miles King , Ken Huang

Post-deployment monitoring of artificial intelligence (AI) systems in health care is essential to ensure their safety, quality, and sustained benefit-and to support governance decisions about which systems to update, modify, or…

Indirect prompt injection attacks threaten AI agents that execute consequential actions, motivating deterministic system-level defenses. Such defenses can provably block unsafe actions by enforcing confidentiality and integrity policies,…

AI coding assistants are now central to professional software development, yet their impact on how developers think about and practice security remains poorly understood. While prior work has documented vulnerability rates in AI-generated…

We implemented and evaluated an automated cyber defense agent. The agent takes security alerts as input and uses reinforcement learning to learn a policy for executing predefined defensive measures. The defender policies were trained in an…

密码学与安全 · 计算机科学 2023-04-24 Jakob Nyberg , Pontus Johnson