中文
相关论文

相关论文: Frontier AI Risk Management Framework in Practice:…

200 篇论文

Human oversight of AI is promoted as a safeguard against risks such as inaccurate outputs, system malfunctions, or violations of fundamental rights, and is mandated in regulation like the European AI Act. Yet debates on human oversight have…

密码学与安全 · 计算机科学 2026-03-06 Jonas C. Ditz , Veronika Lazar , Elmar Lichtmeß , Carola Plesch , Matthias Heck , Kevin Baum , Markus Langer

As foundation models grow in both popularity and capability, researchers have uncovered a variety of ways that the models can pose a risk to the model's owner, user, or others. Despite the efforts of measuring these risks via benchmarks and…

密码学与安全 · 计算机科学 2025-06-04 David Piorkowski , Michael Hind , John Richards , Jacquelyn Martino

This chapter introduces a conceptual framework for qualitative risk assessment of AI, particularly in the context of the EU AI Act. The framework addresses the complexities of legal compliance and fundamental rights protection by itegrating…

Safety cases - clear, assessable arguments for the safety of a system in a given context - are a widely-used technique across various industries for showing a decision-maker (e.g. boards, customers, third parties) that a system is safe. In…

计算机与社会 · 计算机科学 2025-03-10 Benjamin Hilton , Marie Davidsen Buhl , Tomek Korbak , Geoffrey Irving

International institutions may have an important role to play in ensuring advanced AI systems benefit humanity. International collaborations can unlock AI's ability to further sustainable development, and coordination of regulatory efforts…

As intelligence increases, so does its shadow. AI deception, in which systems induce false beliefs to secure self-beneficial outcomes, has evolved from a speculative concern to an empirically demonstrated risk across language models, AI…

Artificial Intelligence (AI) systems introduce unprecedented privacy challenges as they process increasingly sensitive data. Traditional privacy frameworks prove inadequate for AI technologies due to unique characteristics such as…

密码学与安全 · 计算机科学 2025-10-06 Grace Billiris , Asif Gill , Madhushi Bandara

There is an urgent need to identify both short and long-term risks from newly emerging types of Artificial Intelligence (AI), as well as available risk management measures. In response, and to support global efforts in regulating AI and…

计算机与社会 · 计算机科学 2024-11-18 Rokas Gipiškis , Ayrton San Joaquin , Ze Shen Chin , Adrian Regenfuß , Ariel Gil , Koen Holtman

In recent years Artificial Intelligence (AI) has gained much popularity, with the scientific community as well as with the public. AI is often ascribed many positive impacts for different social domains such as medicine and the economy. On…

计算机与社会 · 计算机科学 2021-02-01 Kimon Kieslich , Marco Lünich , Frank Marcinkowski

Powerful new frontier AI technologies are bringing many benefits to society but at the same time bring new risks. AI developers and regulators are therefore seeking ways to assure the safety of such systems, and one promising method under…

计算机与社会 · 计算机科学 2025-02-11 Stephen Barrett , Philip Fox , Joshua Krook , Tuneer Mondal , Simon Mylius , Alejandro Tlaie

As frontier AI systems advance toward transformative capabilities, we need a parallel transformation in how we measure and evaluate these systems to ensure safety and inform governance. While benchmarks have been the primary method for…

人工智能 · 计算机科学 2025-05-12 Markov Grey , Charbel-Raphaël Segerie

Artificial Intelligence (AI) is one of the most transformative technologies of the 21st century. The extent and scope of future AI capabilities remain a key uncertainty, with widespread disagreement on timelines and potential impacts. As…

人工智能 · 计算机科学 2023-11-27 Kyle A. Kilian , Christopher J. Ventura , Mark M. Bailey

Recent advancements in the field of Artificial Intelligence (AI) establish the basis to address challenging tasks. However, with the integration of AI, new risks arise. Therefore, to benefit from its advantages, it is essential to…

机器学习 · 计算机科学 2024-12-20 Ronald Schnitzer , Andreas Hapfelmeier , Sven Gaube , Sonja Zillner

AI evaluations are an important component of the AI governance toolkit, underlying current approaches to safety cases for preventing catastrophic risks. Our paper examines what these evaluations can and cannot tell us. Evaluations can…

计算机与社会 · 计算机科学 2024-12-13 Peter Barnett , Lisa Thiergart

This paper contributes to the nascent debate around safety cases for frontier AI systems. Safety cases are structured, defensible arguments that a system is acceptably safe to deploy in a given context. Historically, they have been used in…

计算机与社会 · 计算机科学 2026-03-11 Shaun Feakins , Ibrahim Habli , Phillip Morgan

AI companies and governments are increasingly concerned about frontier AI systems enabling cybercrime, yet defining meaningful capability thresholds requires knowing the scale of cybercrime today. Current estimates of global cybercrime…

计算机与社会 · 计算机科学 2026-03-24 Kamilė Lukošiūtė , John Halstead , Luca Righetti

Increasingly multi-purpose AI models, such as cutting-edge large language models or other 'general-purpose AI' (GPAI) models, 'foundation models,' generative AI models, and 'frontier models' (typically all referred to hereafter with the…

Artificial intelligence (AI) systems can provide many beneficial capabilities but also risks of adverse events. Some AI systems could present risks of events with very high or catastrophic consequences at societal scale. The US National…

计算机与社会 · 计算机科学 2023-02-24 Anthony M. Barrett , Dan Hendrycks , Jessica Newman , Brandie Nonnecke

An Artificial Intelligence (AI) agent is a software entity that autonomously performs tasks or makes decisions based on pre-defined objectives and data inputs. AI agents, capable of perceiving user inputs, reasoning and planning tasks, and…

密码学与安全 · 计算机科学 2025-11-26 Zehang Deng , Yongjian Guo , Changzhou Han , Wanlun Ma , Junwu Xiong , Sheng Wen , Yang Xiang

Since the publication of the first International AI Safety Report, AI capabilities have continued to improve across key domains. New training techniques that teach AI systems to reason step-by-step and inference-time enhancements have…