中文
相关论文

相关论文: Evaluating Control Protocols for Untrusted AI Agen…

200 篇论文

While artificial intelligence (AI) is advancing rapidly and mastering increasingly complex problems with astonishing performance, the safety assurance of such systems is a major concern. Particularly in the context of safety-critical,…

人工智能 · 计算机科学 2025-07-01 Lars Ullrich , Walter Zimmer , Ross Greer , Knut Graichen , Alois C. Knoll , Mohan Trivedi

This study explores the dynamics of trust in artificial intelligence (AI) agents, particularly large language models (LLMs), by introducing the concept of "deferred trust", a cognitive mechanism where distrust in human agents redirects…

人机交互 · 计算机科学 2025-11-24 Johan Sebastián Galindez-Acosta , Juan José Giraldo-Huertas

How to detect and mitigate deceptive AI systems is an open problem for the field of safe and trustworthy AI. We analyse two algorithms for mitigating deception: The first is based on the path-specific objectives framework where paths in the…

人工智能 · 计算机科学 2023-06-27 Ismail Sahbane , Francis Rhys Ward , C Henrik Åslund

Contemporary benchmarks for agentic artificial intelligence (AI) frequently evaluate safety through isolated task-level accuracy thresholds, implicitly treating autonomous systems as single points of failure. This single-channel paradigm…

计算机与社会 · 计算机科学 2026-02-24 Nelu D. Radpour

Human-AI collaboration has the potential to transform various domains by leveraging the complementary strengths of human experts and Artificial Intelligence (AI) systems. However, unobserved confounding can undermine the effectiveness of…

人机交互 · 计算机科学 2025-02-27 Ruijiang Gao , Mingzhang Yin

Safe multi-agent coordination in uncertain environments can benefit from learning constraints from other agents. Implicitly communicating safety constraints through actions is a promising approach, allowing agents to coordinate and maintain…

系统与控制 · 电气工程与系统科学 2026-04-06 Minh Nguyen , Jingqi Li , Gechen Qu , Claire J. Tomlin

Recent AI systems compress the distance between capability growth and capability deployment. Earlier high-risk technologies were slowed by capital intensity, physical bottlenecks, organizational inertia, and specialized supply chains. By…

人工智能 · 计算机科学 2026-05-05 Wesley Shu , Peng Wei

Automated control monitors could play an important role in overseeing highly capable AI agents that we do not fully trust. Prior work has explored control monitoring in simplified settings, but scaling monitoring to real-world deployments…

Evaluating the safety of AI Systems is a pressing concern for organizations deploying them. In addition to the societal damage done by the lack of fairness of those systems, deployers are concerned about the legal repercussions and the…

Systems operating in adversarial environments may inadvertently leak sensitive information to adversaries. To address this challenge, we revisit the linear-quadratic control framework and introduce deception to actively mislead adversaries.…

最优化与控制 · 数学 2026-04-02 Yerin Kim , Haosheng Zhou , Alexander Benvenuti , Ruimeng Hu , Matthew Hale

AI control protocols use monitors to detect attacks by untrusted AI agents, but standard single-score monitors face two limitations: they miss subtle attacks where outputs look clean but reasoning is off, and they collapse to near-zero…

密码学与安全 · 计算机科学 2026-04-07 Khanh Linh Nguyen , Hoa Nghiem , Tu Tran

A red team simulates adversary attacks to help defenders find effective strategies to defend their systems in a real-world operational setting. As more enterprise systems adopt AI, red-teaming will need to evolve to address the unique…

Oversight and control, which we collectively call supervision, are often discussed as ways to ensure that AI systems are accountable, reliable, and able to fulfill governance and management requirements. However, the requirements for "human…

人工智能 · 计算机科学 2025-11-04 David Manheim , Aidan Homewood

Increased delegation of commercial, scientific, governmental, and personal activities to AI agents -- systems capable of pursuing complex goals with limited supervision -- may exacerbate existing societal risks and introduce new risks.…

Cyber-secure networked control is modeled, analyzed, and experimentally illustrated in this paper. An attack space defined by the adversary's system knowledge, disclosure, and disruption resources is introduced. Adversaries constrained by…

最优化与控制 · 数学 2012-12-04 André Teixeira , Iman Shames , Henrik Sandberg , Karl H. Johansson

AI agents are increasingly deployed across diverse domains to automate complex workflows through long-horizon and high-stakes action executions. Due to their high capability and flexibility, such agents raise significant security and safety…

In the vast domain of cybersecurity, the transition from reactive defense to offensive has become critical in protecting digital infrastructures. This paper explores the integration of Artificial Intelligence (AI) into offensive…

密码学与安全 · 计算机科学 2024-06-13 Leroy Jacob Valencia

The increasing adoption of Reinforcement Learning in safety-critical systems domains such as autonomous vehicles, health, and aviation raises the need for ensuring their safety. Existing safety mechanisms such as adversarial training,…

机器学习 · 计算机科学 2021-11-11 Paulina Stevia Nouwou Mindom , Amin Nikanjam , Foutse Khomh , John Mullins

As LLM agents gain a greater capacity to cause harm, AI developers might increasingly rely on control measures such as monitoring to justify that they are safe. We sketch how developers could construct a "control safety case", which is a…

人工智能 · 计算机科学 2025-01-30 Tomek Korbak , Joshua Clymer , Benjamin Hilton , Buck Shlegeris , Geoffrey Irving

Artificial Intelligence (AI) holds the potential to dramatically improve patient care. However, it is not infallible, necessitating human-AI-collaboration to ensure safe implementation. One aspect of AI safety is the models' ability to…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Anna M. Wundram , Christian F. Baumgartner