中文
相关论文

相关论文: International AI Safety Report 2025: Second Key Up…

200 篇论文

Since the publication of the first International AI Safety Report, AI capabilities have continued to improve across key domains. New training techniques that teach AI systems to reason step-by-step and inference-time enhancements have…

The recent development of powerful AI systems has highlighted the need for robust risk management frameworks in the AI industry. Although companies have begun to implement safety frameworks, current approaches often lack the systematic…

人工智能 · 计算机科学 2025-02-20 Simeon Campos , Henry Papadatos , Fabien Roger , Chloé Touzet , Otter Quarks , Malcolm Murray

Following the AI Seoul Summit in 2024, twelve AI companies published frontier AI safety frameworks (Frameworks) outlining their approaches to managing catastrophic risks from advanced AI systems. Emerging legislation increasingly treats…

计算机与社会 · 计算机科学 2026-05-01 Lily Stelling , Malcolm Murray , Bruno Galizzi , Max Schaffelder , Siméon Campos , Henry Papadatos

International institutions may have an important role to play in ensuring advanced AI systems benefit humanity. International collaborations can unlock AI's ability to further sustainable development, and coordination of regulatory efforts…

Safety has become the central value around which dominant AI governance efforts are being shaped. Recently, this culminated in the publication of the International AI Safety Report, written by 96 experts of which 30 nominated by the…

计算机与社会 · 计算机科学 2025-03-10 Roel Dobbe

To understand and identify the unprecedented risks posed by rapidly advancing artificial intelligence (AI) models, Frontier AI Risk Management Framework in Practice presents a comprehensive assessment of their frontier risks. As Large…

As part of the Frontier AI Safety Commitments agreed to at the 2024 AI Seoul Summit, many AI developers agreed to publish a safety framework outlining how they will manage potential severe risks associated with their systems. This paper…

计算机与社会 · 计算机科学 2025-03-10 Marie Davidsen Buhl , Ben Bucknall , Tammy Masterson

Frontier AI companies first deploy their most advanced models internally, for weeks or months of safety testing, evaluation, and iteration, before a possible public release. For example, Anthropic recently developed a new class of model…

计算机与社会 · 计算机科学 2026-04-30 Oscar Delaney , Sambhav Maheshwari , Joe O'Brien , Theo Bearman , Oliver Guest

As AI systems become more advanced, concerns about large-scale risks from misuse or accidents have grown. This report analyzes the technical research into safe AI development being conducted by three leading AI companies: Anthropic, Google…

计算机与社会 · 计算机科学 2024-09-26 Oscar Delaney , Oliver Guest , Zoe Williams

Over the past year, artificial intelligence (AI) companies have been increasingly adopting AI safety frameworks. These frameworks outline how companies intend to keep the potential risks associated with developing and deploying frontier AI…

计算机与社会 · 计算机科学 2024-09-16 Jide Alaga , Jonas Schuett , Markus Anderljung

Rapidly improving AI capabilities and autonomy hold significant promise of transformation, but are also driving vigorous debate on how to ensure that AI is safe, i.e., trustworthy, reliable, and secure. Building a trusted ecosystem is…

The malicious use or malfunction of advanced general-purpose AI (GPAI) poses risks that, according to leading experts, could lead to the 'marginalisation or extinction of humanity.' To address these risks, there are an increasing number of…

计算机与社会 · 计算机科学 2025-03-26 Rebecca Scholefield , Samuel Martin , Otto Barten

Artificial Intelligence (AI) has rapidly evolved over the past decade and has advanced in areas such as language comprehension, image and video recognition, programming, and scientific reasoning. Recent AI technologies based on large…

机器学习 · 计算机科学 2024-10-30 Jonghong Jeon

The increasing use of AI technologies has led to increasing AI incidents, posing risks and causing harm to individuals, organizations, and society. This study recognizes and addresses the lack of standardized protocols for reliably and…

计算机与社会 · 计算机科学 2025-01-28 Avinash Agarwal , Manisha J Nene

There is an urgent need to identify both short and long-term risks from newly emerging types of Artificial Intelligence (AI), as well as available risk management measures. In response, and to support global efforts in regulating AI and…

计算机与社会 · 计算机科学 2024-11-18 Rokas Gipiškis , Ayrton San Joaquin , Ze Shen Chin , Adrian Regenfuß , Ariel Gil , Koen Holtman

The International AI Safety Report 2026 synthesises the current scientific evidence on the capabilities, emerging risks, and safety of general-purpose AI systems. The report series was mandated by the nations attending the AI Safety Summit…

计算机与社会 · 计算机科学 2026-02-25 Yoshua Bengio , Stephen Clare , Carina Prunkl , Maksym Andriushchenko , Ben Bucknall , Malcolm Murray , Rishi Bommasani , Stephen Casper , Tom Davidson , Raymond Douglas , David Duvenaud , Philip Fox , Usman Gohar , Rose Hadshar , Anson Ho , Tiancheng Hu , Cameron Jones , Sayash Kapoor , Atoosa Kasirzadeh , Sam Manning , Nestor Maslej , Vasilios Mavroudis , Conor McGlynn , Richard Moulange , Jessica Newman , Kwan Yee Ng , Patricia Paskov , Shalaleh Rismani , Girish Sastry , Elizabeth Seger , Scott Singer , Charlotte Stix , Lucia Velasco , Nicole Wheeler , Daron Acemoglu , Vincent Conitzer , Thomas G. Dietterich , Fredrik Heintz , Geoffrey Hinton , Nick Jennings , Susan Leavy , Teresa Ludermir , Vidushi Marda , Helen Margetts , John McDermid , Jane Munga , Arvind Narayanan , Alondra Nelson , Clara Neppel , Sarvapali D. Ramchurn , Stuart Russell , Marietje Schaake , Bernhard Schölkopf , Alvaro Soto , Lee Tiedrich , Gaël Varoquaux , Andrew Yao , Ya-Qin Zhang , Leandro Angelo Aguirre , Olubunmi Ajala , Fahad Albalawi , Noora AlMalek , Christian Busch , Jonathan Collas , André Carlos Ponce de Leon Ferreira de Carvalho , Amandeep Gill , Ahmet Halit Hatip , Juha Heikkilä , Chris Johnson , Gill Jolly , Ziv Katzir , Mary N. Kerema , Hiroaki Kitano , Antonio Krüger , Kyoung Mu Lee , José Ramón López Portillo , Aoife McLysaght , Oleksii Molchanovskyi , Andrea Monti , Mona Nemer , Nuria Oliver , Raquel Pezoa , Audrey Plonk , Balaraman Ravindran , Hammam Riza , Crystal Rugege , Haroon Sheikh , Denise Wong , Yi Zeng , Liming Zhu , Daniel Privitera , Sören Mindermann

As AI models scale to billions of parameters and operate with increasing autonomy, ensuring their safe, reliable operation demands engineering-grade security and assurance frameworks. This paper presents an enterprise-level, risk-aware,…

密码学与安全 · 计算机科学 2025-05-13 Krti Tallam

Rapidly evolving AI exhibits increasingly strong autonomy and goal-directed capabilities, accompanied by derivative systemic risks that are more unpredictable, difficult to control, and potentially irreversible. However, current AI safety…

Increasingly multi-purpose AI models, such as cutting-edge large language models or other 'general-purpose AI' (GPAI) models, 'foundation models,' generative AI models, and 'frontier models' (typically all referred to hereafter with the…

‹ 上一页 1 2 3 10 下一页 ›