English
Related papers

Related papers: Coordinated pausing: An evaluation-based coordinat…

200 papers

The use of Artificial Intelligence (AI) in high-risk, decision-making scenarios presents technical, safety, and normative challenges; problems that may only be ameliorated by human oversight. However, notions of human oversight lack a…

Safety and responsibility evaluations of advanced AI models are a critical but developing field of research and practice. In the development of Google DeepMind's advanced AI models, we innovated on and applied a broad set of approaches to…

Advanced AI systems may be developed which exhibit capabilities that present significant risks to public safety or security. They may also exhibit capabilities that may be applied defensively in a wide set of domains, including (but not…

Computers and Society · Computer Science 2024-10-08 Joe O'Brien , Shaun Ee , Jam Kraprayoon , Bill Anderson-Samways , Oscar Delaney , Zoe Williams

Safety cases - clear, assessable arguments for the safety of a system in a given context - are a widely-used technique across various industries for showing a decision-maker (e.g. boards, customers, third parties) that a system is safe. In…

Computers and Society · Computer Science 2025-03-10 Benjamin Hilton , Marie Davidsen Buhl , Tomek Korbak , Geoffrey Irving

This second update to the 2025 International AI Safety Report assesses new developments in general-purpose AI risk management over the past year. It examines how researchers, public institutions, and AI developers are approaching risk…

Frontier AI models -- highly capable foundation models at the cutting edge of AI development -- may pose severe risks to public safety, human rights, economic stability, and societal value in the coming years. These risks could arise from…

Computers and Society · Computer Science 2025-03-11 Deepika Raman , Nada Madkour , Evan R. Murphy , Krystal Jackson , Jessica Newman

The emergence of pre-trained AI systems with powerful capabilities across a diverse and ever-increasing set of complex domains has raised a critical challenge for AI safety as tasks can become too complicated for humans to judge directly.…

Artificial Intelligence · Computer Science 2023-11-27 Jonah Brown-Cohen , Geoffrey Irving , Georgios Piliouras

A Collaborative Artificial Intelligence System (CAIS) performs actions in collaboration with the human to achieve a common goal. CAISs can use a trained AI model to control human-system interaction, or they can use human interaction to…

Software Engineering · Computer Science 2024-01-24 Diaeddin Rimawi , Antonio Liotta , Marco Todescato , Barbara Russo

As Artificial Intelligence (AI) systems proliferate, the need for systematic, transparent, and actionable processes for evaluating them is growing. While many resources exist to support AI evaluation, they have several limitations. Few…

Computers and Society · Computer Science 2026-02-02 Rachel M. Kim , Blaine Kuehnert , Alice Lai , Kenneth Holstein , Hoda Heidari , Rayid Ghani

Recent proposals for regulating frontier AI models have sparked concerns about the cost of safety regulation, and most such regulations have been shelved due to the safety-innovation tradeoff. This paper argues for an alternative regulatory…

Artificial Intelligence · Computer Science 2025-10-17 Shriyash Upadhyay , Chaithanya Bandi , Narmeen Oozeer , Philip Quirke

Artificial Intelligence (AI) Safety Institutes and governments worldwide are deciding whether they evaluate advanced AI themselves, support a private evaluation ecosystem or do both. Evaluation regimes have been established in a wide range…

Computers and Society · Computer Science 2025-08-06 Merlin Stein , Milan Gandhi , Theresa Kriecherbauer , Amin Oueslati , Robert Trager

AI safety has emerged as a critical priority as these systems are increasingly deployed in real-world applications. We propose the first domain-agnostic AI safety ensuring framework that achieves strong safety guarantees while preserving…

Artificial Intelligence · Computer Science 2025-10-07 Beomjun Kim , Kangyeon Kim , Sunwoo Kim , Yeonsang Shin , Heejin Ahn

Collision detection via visual fences can significantly enhance the safety of collaborative robotic arms. Existing work typically performs such detection based on pre-deployed stationary cameras outside the robotic arm's workspace. These…

Robotics · Computer Science 2024-03-12 Xian Huang , Yuanjiong Ying , Wei Dong

Open-weight advanced AI models -- systems whose parameters are freely available for download and adaptation -- are reshaping the global AI landscape. As these models rapidly close the performance gap with closed alternatives, they enable…

Computers and Society · Computer Science 2026-02-24 Bengüsu Özcan , Alex Petropoulos , Max Reddel

We examine how the federal government can enhance its AI emergency preparedness: the ability to detect and prepare for time-sensitive national security threats relating to AI. Emergency preparedness can improve the government's ability to…

Computers and Society · Computer Science 2024-07-30 Akash Wasil , Everett Smith , Corin Katzke , Justin Bullock

There is still a significant gap between expectations and the successful adoption of AI to innovate and improve businesses. Due to the emergence of deep learning, AI adoption is more complex as it often incorporates big data and the…

Artificial Intelligence · Computer Science 2022-09-16 Dian Tjondronegoro , Elizabeth Yuwono , Brent Richards , Damian Green , Siiri Hatakka

As AI rapidly advances, the security risks posed by AI are becoming increasingly severe, especially in critical scenarios, including those posing existential risks. If AI becomes uncontrollable, manipulated, or actively evades safety…

Artificial Intelligence · Computer Science 2025-08-29 Donglin Wang , Weiyun Liang , Chunyuan Chen , Jing Xu , Yulong Fu

The April 2026 disclosure that a frontier large language model escaped its security sandbox, executed unauthorized actions, and concealed its modifications to version control history demonstrates that agentic AI systems with autonomous tool…

Cryptography and Security · Computer Science 2026-04-28 Richard Joseph Mitchell

In human-AI decision making, designing AI that complements human expertise has been a natural strategy to enhance human-AI collaboration, yet it often comes at the cost of decreased AI performance in areas of human strengths. This can…

Artificial Intelligence · Computer Science 2026-02-24 Hasan Amin , Ming Yin , Rajiv Khanna

Is there a way to design powerful AI systems based on machine learning methods that would satisfy probabilistic safety guarantees? With the long-term goal of obtaining a probabilistic guarantee that would apply in every context, we consider…

Artificial Intelligence · Computer Science 2025-06-17 Yoshua Bengio , Michael K. Cohen , Nikolay Malkin , Matt MacDermott , Damiano Fornasiere , Pietro Greiner , Younesse Kaddar