中文
相关论文

相关论文: The Singapore Consensus on Global AI Safety Resear…

200 篇论文

Our survey of 53 specialists across 105 AI reliability and security research areas identifies the most promising research prospects to guide strategic AI R&D investment. As companies are seeking to develop AI systems with broadly…

计算机与社会 · 计算机科学 2025-05-29 Joe O'Brien , Jeremy Dolan , Jay Kim , Jonah Dykhuizen , Jeba Sania , Sebastian Becker , Jam Kraprayoon , Cara Labrador

Safety has become the central value around which dominant AI governance efforts are being shaped. Recently, this culminated in the publication of the International AI Safety Report, written by 96 experts of which 30 nominated by the…

计算机与社会 · 计算机科学 2025-03-10 Roel Dobbe

The International AI Safety Report 2026 synthesises the current scientific evidence on the capabilities, emerging risks, and safety of general-purpose AI systems. The report series was mandated by the nations attending the AI Safety Summit…

计算机与社会 · 计算机科学 2026-02-25 Yoshua Bengio , Stephen Clare , Carina Prunkl , Maksym Andriushchenko , Ben Bucknall , Malcolm Murray , Rishi Bommasani , Stephen Casper , Tom Davidson , Raymond Douglas , David Duvenaud , Philip Fox , Usman Gohar , Rose Hadshar , Anson Ho , Tiancheng Hu , Cameron Jones , Sayash Kapoor , Atoosa Kasirzadeh , Sam Manning , Nestor Maslej , Vasilios Mavroudis , Conor McGlynn , Richard Moulange , Jessica Newman , Kwan Yee Ng , Patricia Paskov , Shalaleh Rismani , Girish Sastry , Elizabeth Seger , Scott Singer , Charlotte Stix , Lucia Velasco , Nicole Wheeler , Daron Acemoglu , Vincent Conitzer , Thomas G. Dietterich , Fredrik Heintz , Geoffrey Hinton , Nick Jennings , Susan Leavy , Teresa Ludermir , Vidushi Marda , Helen Margetts , John McDermid , Jane Munga , Arvind Narayanan , Alondra Nelson , Clara Neppel , Sarvapali D. Ramchurn , Stuart Russell , Marietje Schaake , Bernhard Schölkopf , Alvaro Soto , Lee Tiedrich , Gaël Varoquaux , Andrew Yao , Ya-Qin Zhang , Leandro Angelo Aguirre , Olubunmi Ajala , Fahad Albalawi , Noora AlMalek , Christian Busch , Jonathan Collas , André Carlos Ponce de Leon Ferreira de Carvalho , Amandeep Gill , Ahmet Halit Hatip , Juha Heikkilä , Chris Johnson , Gill Jolly , Ziv Katzir , Mary N. Kerema , Hiroaki Kitano , Antonio Krüger , Kyoung Mu Lee , José Ramón López Portillo , Aoife McLysaght , Oleksii Molchanovskyi , Andrea Monti , Mona Nemer , Nuria Oliver , Raquel Pezoa , Audrey Plonk , Balaraman Ravindran , Hammam Riza , Crystal Rugege , Haroon Sheikh , Denise Wong , Yi Zeng , Liming Zhu , Daniel Privitera , Sören Mindermann

Since the publication of the first International AI Safety Report, AI capabilities have continued to improve across key domains. New training techniques that teach AI systems to reason step-by-step and inference-time enhancements have…

This second update to the 2025 International AI Safety Report assesses new developments in general-purpose AI risk management over the past year. It examines how researchers, public institutions, and AI developers are approaching risk…

The first International AI Safety Report comprehensively synthesizes the current evidence on the capabilities, risks, and safety of advanced AI systems. The report was mandated by the nations attending the AI Safety Summit in Bletchley, UK.…

计算机与社会 · 计算机科学 2025-01-30 Yoshua Bengio , Sören Mindermann , Daniel Privitera , Tamay Besiroglu , Rishi Bommasani , Stephen Casper , Yejin Choi , Philip Fox , Ben Garfinkel , Danielle Goldfarb , Hoda Heidari , Anson Ho , Sayash Kapoor , Leila Khalatbari , Shayne Longpre , Sam Manning , Vasilios Mavroudis , Mantas Mazeika , Julian Michael , Jessica Newman , Kwan Yee Ng , Chinasa T. Okolo , Deborah Raji , Girish Sastry , Elizabeth Seger , Theodora Skeadas , Tobin South , Emma Strubell , Florian Tramèr , Lucia Velasco , Nicole Wheeler , Daron Acemoglu , Olubayo Adekanmbi , David Dalrymple , Thomas G. Dietterich , Edward W. Felten , Pascale Fung , Pierre-Olivier Gourinchas , Fredrik Heintz , Geoffrey Hinton , Nick Jennings , Andreas Krause , Susan Leavy , Percy Liang , Teresa Ludermir , Vidushi Marda , Helen Margetts , John McDermid , Jane Munga , Arvind Narayanan , Alondra Nelson , Clara Neppel , Alice Oh , Gopal Ramchurn , Stuart Russell , Marietje Schaake , Bernhard Schölkopf , Dawn Song , Alvaro Soto , Lee Tiedrich , Gaël Varoquaux , Andrew Yao , Ya-Qin Zhang , Fahad Albalawi , Marwan Alserkal , Olubunmi Ajala , Guillaume Avrin , Christian Busch , André Carlos Ponce de Leon Ferreira de Carvalho , Bronwyn Fox , Amandeep Singh Gill , Ahmet Halit Hatip , Juha Heikkilä , Gill Jolly , Ziv Katzir , Hiroaki Kitano , Antonio Krüger , Chris Johnson , Saif M. Khan , Kyoung Mu Lee , Dominic Vincent Ligot , Oleksii Molchanovskyi , Andrea Monti , Nusu Mwamanzi , Mona Nemer , Nuria Oliver , José Ramón López Portillo , Balaraman Ravindran , Raquel Pezoa Rivera , Hammam Riza , Crystal Rugege , Ciarán Seoighe , Jerry Sheehan , Haroon Sheikh , Denise Wong , Yi Zeng

Ensuring that AI systems reliably and robustly avoid harmful or dangerous behaviours is a crucial challenge, especially for AI systems with a high degree of autonomy and general intelligence, or systems used in safety-critical contexts. In…

This policy report draws on country studies from China, South Korea, Singapore, and the United Kingdom to identify effective tools and key barriers to interoperability in AI safety governance. It offers practical recommendations to support…

计算机与社会 · 计算机科学 2026-01-13 Yik Chan Chin , David A. Raho , Hag-Min Kim , Chunli Bi , James Ong , Jingbo Huang , Serge Stinckwich

Recent advances in machine learning, particularly the emergence of foundation models, are leading to new opportunities to develop technology-based solutions to societal problems. However, the reasoning and inner workings of today's complex…

计算机与社会 · 计算机科学 2025-07-01 Rajeev Alur , Greg Durrett , Hadas Kress-Gazit , Corina Păsăreanu , René Vidal

Artificial intelligence (AI) technologies (re-)shape modern life, driving innovation in a wide range of sectors. However, some AI systems have yielded unexpected or undesirable outcomes or have been used in questionable manners. As a…

Last decade has seen major improvements in the performance of artificial intelligence which has driven wide-spread applications. Unforeseen effects of such mass-adoption has put the notion of AI safety into the public eye. AI safety is a…

计算机与社会 · 计算机科学 2020-07-10 Mislav Juric , Agneza Sandic , Mario Brcic

To facilitate the widespread acceptance of AI systems guiding decision-making in real-world applications, it is key that solutions comprise trustworthy, integrated human-AI systems. Not only in safety-critical applications such as…

人工智能 · 计算机科学 2020-01-16 Florian Buettner , John Piorkowski , Ian McCulloh , Ulli Waltinger

The conversation around artificial intelligence (AI) often focuses on safety, transparency, accountability, alignment, and responsibility. However, AI security (i.e., the safeguarding of data, models, and pipelines from adversarial…

密码学与安全 · 计算机科学 2025-04-24 Krti Tallam

While artificial intelligence (AI) is advancing rapidly and mastering increasingly complex problems with astonishing performance, the safety assurance of such systems is a major concern. Particularly in the context of safety-critical,…

人工智能 · 计算机科学 2025-07-01 Lars Ullrich , Walter Zimmer , Ross Greer , Knut Graichen , Alois C. Knoll , Mohan Trivedi

The increasing use of AI technologies has led to increasing AI incidents, posing risks and causing harm to individuals, organizations, and society. This study recognizes and addresses the lack of standardized protocols for reliably and…

计算机与社会 · 计算机科学 2025-01-28 Avinash Agarwal , Manisha J Nene

AI safety benchmarks are pivotal for safety in advanced AI systems; however, they have significant technical, epistemic, and sociotechnical shortcomings. We present a review of 210 safety benchmarks that maps out common challenges in safety…

计算机与社会 · 计算机科学 2026-02-10 Cheng Yu , Severin Engelmann , Ruoxuan Cao , Dalia Ali , Orestis Papakyriakopoulos

Recent discussions and research in AI safety have increasingly emphasized the deep connection between AI safety and existential risk from advanced AI systems, suggesting that work on AI safety necessarily entails serious consideration of…

计算机与社会 · 计算机科学 2025-02-17 Balint Gyevnar , Atoosa Kasirzadeh

As AI systems become more advanced, concerns about large-scale risks from misuse or accidents have grown. This report analyzes the technical research into safe AI development being conducted by three leading AI companies: Anthropic, Google…

计算机与社会 · 计算机科学 2024-09-26 Oscar Delaney , Oliver Guest , Zoe Williams

This paper contributes to the nascent debate around safety cases for frontier AI systems. Safety cases are structured, defensible arguments that a system is acceptably safe to deploy in a given context. Historically, they have been used in…

计算机与社会 · 计算机科学 2026-03-11 Shaun Feakins , Ibrahim Habli , Phillip Morgan
‹ 上一页 1 2 3 10 下一页 ›