中文
相关论文

相关论文: Risk thresholds for frontier AI

200 篇论文

Collaborative AI systems (CAISs) aim at working together with humans in a shared space to achieve a common goal. This critical setting yields hazardous circumstances that could harm human beings. Thus, building such systems with strong…

Recent concern about harms of information technologies motivate consideration of regulatory action to forestall or constrain certain developments in the field of artificial intelligence (AI). However, definitional ambiguity hampers the…

计算机与社会 · 计算机科学 2019-12-25 P. M. Krafft , Meg Young , Michael Katell , Karen Huang , Ghislain Bugingo

In this work, we present and analyze reported failures of artificially intelligent systems and extrapolate our analysis to future AIs. We suggest that both the frequency and the seriousness of future AI failures will steadily increase. AI…

人工智能 · 计算机科学 2016-10-26 Roman V. Yampolskiy , M. S. Spellchecker

Recent work has demonstrated the plausibility of frontier AI models scheming -- knowingly and covertly pursuing an objective misaligned with its developer's intentions. Such behavior could be very hard to detect, and if present in future…

With the turmoil in cybersecurity and the mind-blowing advances in AI, it is only natural that cybersecurity practitioners consider further employing learning techniques to help secure their organizations and improve the efficiency of their…

密码学与安全 · 计算机科学 2019-12-17 Ricardo Morla

Ensuring responsible use of artificial intelligence (AI) has become imperative as autonomous systems increasingly influence critical societal domains. However, the concept of trustworthy AI remains broad and multi-faceted. This thesis…

人工智能 · 计算机科学 2025-10-28 Filip Cano

Potential advancements in artificial intelligence (AI) could have profound implications for how countries research and develop weapons systems, and how militaries deploy those systems on the battlefield. The idea of AI-enabled military…

计算机与社会 · 计算机科学 2022-11-02 Paul Scharre , Megan Lamberth

Existing strategies for managing risks from advanced AI systems often focus on affecting what AI systems are developed and how they diffuse. However, this approach becomes less feasible as the number of developers of advanced AI grows, and…

计算机与社会 · 计算机科学 2025-01-24 Jamie Bernardi , Gabriel Mukobi , Hilary Greaves , Lennart Heim , Markus Anderljung

This paper examines the state of affairs on Frontier Safety Policies in light of capability progress and growing expectations held by government actors and AI safety researchers from these safety policies. It subsequently argues that FSPs…

计算机与社会 · 计算机科学 2025-01-29 Matteo Pistillo

This paper aims to provide an overview of the ethical concerns in artificial intelligence (AI) and the framework that is needed to mitigate those risks, and to suggest a practical path to ensure the development and use of AI at the United…

计算机与社会 · 计算机科学 2021-04-27 Lambert Hogenhout

Legislation and public sentiment throughout the world have promoted fairness metrics, explainability, and interpretability as prescriptions for the responsible development of ethical artificial intelligence systems. Despite the importance…

人工智能 · 计算机科学 2022-03-08 Erick Galinkin

Advances in AI are widely understood to have implications for cybersecurity. Articles have emphasized the effect of AI on the cyber offense-defense balance, and commentators can be found arguing either that cyber will privilege attackers or…

密码学与安全 · 计算机科学 2025-08-25 Benjamin Murphy , Twm Stone

Data is essential to train and fine-tune today's frontier artificial intelligence (AI) models and to develop future ones. To date, academic, legal, and regulatory work has primarily addressed how data can directly harm consumers and…

人工智能 · 计算机科学 2025-06-03 Jason Hausenloy , Duncan McClements , Madhavendra Thakur

We assess whether AI systems can credibly evaluate investment risk appetite-a task that must be thoroughly validated before automation. Our analysis was conducted on proprietary systems (GPT, Claude, Gemini) and open-weight models (LLaMA,…

The development of Artificial Intelligence (AI), including AI in Science (AIS), should be done following the principles of responsible AI. Progress in responsible AI is often quantified through evaluation metrics, yet there has been less…

计算机与社会 · 计算机科学 2025-10-31 Theresia Veronika Rampisela , Maria Maistro , Tuukka Ruotsalo , Christina Lioma

This paper provides an overview and critique of the risk based model of artificial intelligence (AI) governance that has become a popular approach to AI regulation across multiple jurisdictions. The 'AI Policy Landscape in Europe, North…

计算机与社会 · 计算机科学 2025-07-22 Veve Fry

Embedding artificial intelligence into systems introduces significant challenges to modern engineering practices. Hazard analysis tools and processes have not yet been adequately adapted to the new paradigm. This paper describes initial…

软件工程 · 计算机科学 2022-03-30 Nikolas Martelaro , Carol J. Smith , Tamara Zilovic

This paper provides policy recommendations to reduce extinction risks from advanced artificial intelligence (AI). First, we briefly provide background information about extinction risks from AI. Second, we argue that voluntary commitments…

人工智能 · 计算机科学 2025-09-30 Andrea Miotti

Recent developments in artificial intelligence (AI) have permeated through an array of different immersive environments, including virtual, augmented, and mixed realities. AI brings a wealth of potential that centers on its ability to…

人机交互 · 计算机科学 2024-05-10 Wangfan Li , Rohit Mallick , Carlos Toxtli-Hernandez , Christopher Flathmann , Nathan J. McNeese

Policymakers increasingly use development cost and compute as proxies for AI capabilities and risks. Recent laws have introduced regulatory requirements for models or developers that are contingent on specific thresholds. However, technical…

人工智能 · 计算机科学 2025-07-22 Stephen Casper , Luke Bailey , Tim Schreier