English
Related papers

Related papers: Open Problems in Frontier AI Risk Management

200 papers

Rapid advancements in artificial intelligence (AI) have sparked growing concerns among experts, policymakers, and world leaders regarding the potential for increasingly advanced AI systems to pose catastrophic risks. Although numerous risks…

Computers and Society · Computer Science 2023-10-11 Dan Hendrycks , Mantas Mazeika , Thomas Woodside

Purpose: The governance of artificial iintelligence (AI) systems requires a structured approach that connects high-level regulatory principles with practical implementation. Existing frameworks lack clarity on how regulations translate into…

Computers and Society · Computer Science 2025-09-16 Avinash Agarwal , Manisha J. Nene

In this position paper, we address the persistent gap between rapidly growing AI capabilities and lagging safety progress. Existing paradigms divide into ``Make AI Safe'', which applies post-hoc alignment and guardrails but remains brittle…

Machine Learning · Computer Science 2025-09-09 Youbang Sun , Xiang Wang , Jie Fu , Chaochao Lu , Bowen Zhou

As generative AI systems, including large language models (LLMs) and diffusion models, advance rapidly, their growing adoption has led to new and complex security risks often overlooked in traditional AI risk assessment frameworks. This…

Cryptography and Security · Computer Science 2024-10-21 Aviral Srivastava , Sourav Panda

Problems of cooperation--in which agents seek ways to jointly improve their welfare--are ubiquitous and important. They can be found at scales ranging from our daily routines--such as driving on highways, scheduling meetings, and working…

Artificial Intelligence · Computer Science 2020-12-17 Allan Dafoe , Edward Hughes , Yoram Bachrach , Tantum Collins , Kevin R. McKee , Joel Z. Leibo , Kate Larson , Thore Graepel

We sketch how developers of frontier AI systems could construct a structured rationale -- a 'safety case' -- that an AI system is unlikely to cause catastrophic outcomes through scheming. Scheming is a potential threat model where AI…

Safety frameworks represent a significant development in AI governance: they are the first type of publicly shared catastrophic risk management framework developed by major AI companies and focus specifically on AI scaling decisions. I…

Computers and Society · Computer Science 2024-10-02 Atoosa Kasirzadeh

Organizations and governments that develop, deploy, use, and govern AI must coordinate on effective risk mitigation. However, the landscape of AI risk mitigation frameworks is fragmented, uses inconsistent terminology, and has gaps in…

Computers and Society · Computer Science 2025-12-16 Alexander K. Saeri , Sophia Lloyd George , Jess Graham , Clelia D. Lacarriere , Peter Slattery , Michael Noetel , Neil Thompson

Humanity is progressing towards automated product development, a trend that promises faster creation of better products and thus the acceleration of technological progress. However, increasing reliance on non-human agents for this process…

Computers and Society · Computer Science 2025-06-03 Jan Göpfert , Jann M. Weinand , Patrick Kuckertz , Noah Pflugradt , Jochen Linßen

As artificial intelligence (AI) technologies increasingly enter important sectors like healthcare, transportation, and finance, the development of effective governance frameworks is crucial for dealing with ethical, security, and societal…

Computers and Society · Computer Science 2025-03-11 Amir Al-Maamari

This policy report draws on country studies from China, South Korea, Singapore, and the United Kingdom to identify effective tools and key barriers to interoperability in AI safety governance. It offers practical recommendations to support…

Computers and Society · Computer Science 2026-01-13 Yik Chan Chin , David A. Raho , Hag-Min Kim , Chunli Bi , James Ong , Jingbo Huang , Serge Stinckwich

As AI models scale to billions of parameters and operate with increasing autonomy, ensuring their safe, reliable operation demands engineering-grade security and assurance frameworks. This paper presents an enterprise-level, risk-aware,…

Cryptography and Security · Computer Science 2025-05-13 Krti Tallam

Traditional safety engineering assesses systems in their context of use, e.g. the operational design domain (road layout, speed limits, weather, etc.) for self-driving vehicles (including those using AI). We refer to this as downstream…

Computers and Society · Computer Science 2025-01-13 John McDermid , Yan Jia , Ibrahim Habli

Despite advances in large language model capabilities in recent years, a large gap remains in their capabilities and safety performance for many languages beyond a relatively small handful of globally dominant languages. This paper provides…

As Generative Artificial Intelligence (GenAI) technologies evolve at an unprecedented rate, global governance approaches struggle to keep pace with the technology, highlighting a critical issue in the governance adaptation of significant…

Computers and Society · Computer Science 2024-11-22 Jose Luna , Ivan Tan , Xiaofei Xie , Lingxiao Jiang

As foundation models grow in both popularity and capability, researchers have uncovered a variety of ways that the models can pose a risk to the model's owner, user, or others. Despite the efforts of measuring these risks via benchmarks and…

Cryptography and Security · Computer Science 2025-06-04 David Piorkowski , Michael Hind , John Richards , Jacquelyn Martino

As artificial intelligence transforms a wide range of sectors and drives innovation, it also introduces complex challenges concerning ethics, transparency, bias, and fairness. The imperative for integrating Responsible AI (RAI) principles…

Computers and Society · Computer Science 2024-01-23 Amna Batool , Didar Zowghi , Muneera Bano

As AI technologies increase in capability and ubiquity, AI accidents are becoming more common. Based on normal accident theory, high reliability theory, and open systems theory, we create a framework for understanding the risks associated…

Computers and Society · Computer Science 2024-03-13 Heather M. Williams , Roman V. Yampolskiy

Alignment of artificial intelligence (AI) encompasses the normative problem of specifying how AI systems should act and the technical problem of ensuring AI systems comply with those specifications. To date, AI alignment has generally…

Advanced reasoning models with agentic capabilities (AI agents) are deployed to interact with humans and to solve sequential decision-making problems under (approximate) utility functions and internal models. When such problems have…

Artificial Intelligence · Computer Science 2025-09-25 Daniel Jarne Ornia , Nicholas Bishop , Joel Dyer , Wei-Chen Lee , Ani Calinescu , Doyne Farmer , Michael Wooldridge