中文
相关论文

相关论文: An Example Safety Case for Safeguards Against Misu…

200 篇论文

Adversarial attacks pose a severe risk to AI systems used in healthcare, capable of misleading models into dangerous misclassifications that can delay treatments or cause misdiagnoses. These attacks, often imperceptible to human perception,…

机器学习 · 计算机科学 2025-10-29 Alyssa Gerhart , Balaji Iyangar

Artificial intelligence (AI) is increasingly being used to augment and automate cyber operations, altering the scale, speed, and accessibility of malicious activity. These shifts raise urgent questions about when AI systems introduce…

密码学与安全 · 计算机科学 2026-01-27 Krystal Jackson , Deepika Raman , Jessica Newman , Nada Madkour , Charlotte Yuan , Evan R. Murphy

Control evaluations measure whether monitoring and security protocols for AI systems prevent intentionally subversive AI models from causing harm. Our work presents the first control evaluation performed in an agent environment. We…

机器学习 · 计算机科学 2025-04-15 Aryan Bhatt , Cody Rushing , Adam Kaufman , Tyler Tracy , Vasil Georgiev , David Matolcsi , Akbir Khan , Buck Shlegeris

Recent advances in AI agents capable of solving complex, everyday tasks, from scheduling to customer service, have enabled deployment in real-world settings, but their possibilities for unsafe behavior demands rigorous evaluation. While…

As LLM agents grow more capable of causing harm autonomously, AI developers will rely on increasingly sophisticated control measures to prevent possibly misaligned agents from causing harm. AI developers could demonstrate that their control…

人工智能 · 计算机科学 2025-04-08 Tomek Korbak , Mikita Balesni , Buck Shlegeris , Geoffrey Irving

Several different approaches exist for ensuring the safety of future Transformative Artificial Intelligence (TAI) or Artificial Superintelligence (ASI) systems, and proponents of different approaches have made different and debated claims…

人工智能 · 计算机科学 2022-01-11 Issa Rice , David Manheim

AI scientists powered by large language models have demonstrated substantial promise in autonomously conducting experiments and facilitating scientific discoveries across various disciplines. While their capabilities are promising, these…

The growing integration of AI into cybersecurity is reshaping the balance between attackers and defenders. When access to advanced AI-enabled defence tools is uneven, resource-limited defenders may be unable to adopt effective protection,…

人工智能 · 计算机科学 2026-05-14 Adeela Bashir , Zia Ush Shamszaman , Zhao Song , Matjaz Perc , The Anh Han

Prominent AI companies are producing 'safety frameworks' as a type of voluntary self-governance. These statements purport to establish risk thresholds and safety procedures for the development and deployment of highly capable AI.…

计算机与社会 · 计算机科学 2025-10-14 Sam Coggins , Alexander K. Saeri , Katherine A. Daniell , Lorenn P. Ruster , Jessie Liu , Jenny L. Davis

Developers of some safety critical systems construct a safety case. Developers changing a system during development or after release must analyse the change's impact on the safety case. Evidence might be invalidated by changes to the system…

软件工程 · 计算机科学 2014-04-29 Omar Jaradat , Patrick Graydon , Iain Bate

As AI models are increasingly deployed across diverse real-world scenarios, ensuring their safety remains a critical yet underexplored challenge. While substantial efforts have been made to evaluate and enhance AI safety, the lack of a…

Recent advances in AI are transforming AI's ubiquitous presence in our world from that of standalone AI-applications into deeply integrated AI-agents. These changes have been driven by agents' increasing capability to autonomously make…

密码学与安全 · 计算机科学 2025-07-04 Jose Sanchez Vicarte , Marcin Spoczynski , Mostafa Elsaid

How to make artificial intelligence (AI) systems safe and aligned with human values is an open research question. Proposed solutions tend toward relying on human intervention in uncertain situations, learning human values and intentions…

计算机与社会 · 计算机科学 2023-10-04 Jeffrey W. Johnston

As artificial intelligence (AI) systems become increasingly embedded in critical societal functions, the need for robust red teaming methodologies continues to grow. In this forum piece, we examine emerging approaches to automating AI red…

计算机与社会 · 计算机科学 2025-03-31 Alice Qian Zhang , Jina Suh , Mary L. Gray , Hong Shen

As frontier AI systems advance toward transformative capabilities, we need a parallel transformation in how we measure and evaluate these systems to ensure safety and inform governance. While benchmarks have been the primary method for…

人工智能 · 计算机科学 2025-05-12 Markov Grey , Charbel-Raphaël Segerie

This paper argues that a range of current AI systems have learned how to deceive humans. We define deception as the systematic inducement of false beliefs in the pursuit of some outcome other than the truth. We first survey empirical…

计算机与社会 · 计算机科学 2023-08-29 Peter S. Park , Simon Goldstein , Aidan O'Gara , Michael Chen , Dan Hendrycks

As Automated Driving Systems (ADS) technology advances, ensuring safety and public trust requires robust assurance frameworks, with safety cases emerging as a critical tool toward such a goal. This paper explores an approach to assess how a…

软件工程 · 计算机科学 2025-06-12 Scott Schnelle , Francesca Favaro , Laura Fraade-Blanar , David Wichner , Holland Broce , Justin Miranda

Artificial intelligence (AI) is interacting with people at an unprecedented scale, offering new avenues for immense positive impact, but also raising widespread concerns around the potential for individual and societal harm. Today, the…

人工智能 · 计算机科学 2024-06-25 Andrea Bajcsy , Jaime F. Fisac

The rapid growth of Large Language Models (LLMs) presents significant privacy, security, and ethical concerns. While much research has proposed methods for defending LLM systems against misuse by malicious actors, researchers have recently…

The rapid development of AI systems poses unprecedented risks, including loss of control, misuse, geopolitical instability, and concentration of power. To navigate these risks and avoid worst-case outcomes, governments may proactively…

人工智能 · 计算机科学 2025-07-15 Peter Barnett , Aaron Scher , David Abecassis