中文
相关论文

相关论文: Provably safe systems: the only path to controllab…

200 篇论文

International agreements about AI development may be required to reduce catastrophic risks from advanced AI systems. However, agreements about such a high-stakes technology must be backed by verification mechanisms--processes or tools that…

计算机与社会 · 计算机科学 2025-06-23 Aaron Scher , Lisa Thiergart

A cautious interpretation of AI regulations and policy in the EU and the USA place explainability as a central deliverable of compliant AI systems. However, from a technical perspective, explainable AI (XAI) remains an elusive and complex…

计算机与社会 · 计算机科学 2024-06-14 Neo Christopher Chung , Hongkyou Chung , Hearim Lee , Lennart Brocki , Hongbeom Chung , George Dyer

The recent advancement in artificial intelligence (AI) technologies facilitates a paradigm shift toward automation. Autonomous systems are fully or partially replacing manually crafted ones. At the core of these systems is automated…

人工智能 · 计算机科学 2026-04-14 Mir Md Sajid Sarwar

Ensuring Artificial General Intelligence (AGI) reliably avoids harmful behaviors is a critical challenge, especially for systems with high autonomy or in safety-critical domains. Despite various safety assurance proposals and extreme risk…

计算机与社会 · 计算机科学 2024-12-24 Chao Yang , Chaochao Lu , Yingchun Wang , Bowen Zhou

Explainable Artificial Intelligence (XAI) techniques are frequently required by users in many AI systems with the goal of understanding complex models, their associated predictions, and gaining trust. While suitable for some specific tasks…

人机交互 · 计算机科学 2023-03-22 Savio Rozario , George Čevora

Recent AI progress has outpaced expectations, with some experts now predicting AI that matches or exceeds human capabilities in all cognitive areas (AGI) could emerge this decade, potentially posing grave national and global security…

计算机与社会 · 计算机科学 2025-07-30 Sarah Hastings-Woodhouse

Artificial Intelligence (AI) is a double-edged sword: on one hand, AI promises to provide great advances that could benefit humanity, but on the other hand, AI poses substantial (even existential) risks. With advancements happening daily,…

计算机与社会 · 计算机科学 2024-02-05 Willem van der Maden , Derek Lomas , Malak Sadek , Paul Hekkert

The malicious use or malfunction of advanced general-purpose AI (GPAI) poses risks that, according to leading experts, could lead to the 'marginalisation or extinction of humanity.' To address these risks, there are an increasing number of…

计算机与社会 · 计算机科学 2025-03-26 Rebecca Scholefield , Samuel Martin , Otto Barten

Artificial intelligence is already being applied in and impacting many important sectors in society, including healthcare, finance, and policing. These applications will increase as AI capabilities continue to progress, which has the…

计算机与社会 · 计算机科学 2022-06-23 Jess Whittlestone , Sam Clarke

The young field of AI Safety is still in the process of identifying its challenges and limitations. In this paper, we formally describe one such impossibility result, namely Unpredictability of AI. We prove that it is impossible to…

人工智能 · 计算机科学 2019-05-31 Roman V. Yampolskiy

AI agents have been boosted by large language models. AI agents can function as intelligent assistants and complete tasks on behalf of their users with access to tools and the ability to execute commands in their environments. Through…

密码学与安全 · 计算机科学 2024-12-19 Yifeng He , Ethan Wang , Yuyang Rong , Zifei Cheng , Hao Chen

Artificial Intelligence (AI) and the regulation thereof is a topic that is increasingly being discussed within various fora. Various proposals have been made in literature for defining regulatory bodies and/or related regulation. In this…

计算机与社会 · 计算机科学 2021-08-23 Joshua Ellul , Stephen McCarthy , Trevor Sammut , Juanita Brockdorff , Matthew Scerri , Gordon J. Pace

Artificial intelligence (AI) has been clearly established as a technology with the potential to revolutionize fields from healthcare to finance - if developed and deployed responsibly. This is the topic of responsible AI, which emphasizes…

人工智能 · 计算机科学 2023-12-05 Stephanie Baker , Wei Xiang

Currently, we are in an environment where the fraction of automated vehicles is negligibly small. We anticipate that this fraction will increase in coming decades before if ever, we have a fully automated transportation system. Motivated by…

系统与控制 · 计算机科学 2018-10-16 Xi Liu , Ke Ma , P. R. Kumar

Corrigibility is a safety property for artificially intelligent agents. A corrigible agent will not resist attempts by authorized parties to alter the goals and constraints that were encoded in the agent when it was first started. This…

人工智能 · 计算机科学 2020-04-06 Koen Holtman

With Artificial Intelligence (AI) becoming ubiquitous in every application domain, the need for explanations is paramount to enhance transparency and trust among non-technical users. Despite the potential shown by Explainable AI (XAI) for…

人机交互 · 计算机科学 2024-02-05 Aditya Bhattacharya

Everyone from AI executives and researchers to doomsayers, politicians, and activists is talking about Artificial General Intelligence (AGI). Yet, they often don't seem to agree on its exact definition. One common definition of AGI is an AI…

人工智能 · 计算机科学 2026-03-02 Judah Goldfeder , Philippe Wyder , Yann LeCun , Ravid Shwartz Ziv

This document focuses on the threats, especially near-term threats, that Artificial Intelligence (AI) brings to society. Most of the threats discussed here can result from any algorithmic process, not just AI; in addition, defining AI is…

计算机与社会 · 计算机科学 2024-09-10 Don Byrd

Existing theoretical universal algorithmic intelligence models are not practically realizable. More pragmatic approach to artificial general intelligence is based on cognitive architectures, which are, however, non-universal in sense that…

人工智能 · 计算机科学 2012-09-20 Alexey Potapov , Sergey Rodionov , Andrew Myasnikov , Galymzhan Begimov

Artificial intelligence (AI) technologies (re-)shape modern life, driving innovation in a wide range of sectors. However, some AI systems have yielded unexpected or undesirable outcomes or have been used in questionable manners. As a…