中文
相关论文

相关论文: Principles for new ASI Safety Paradigms

200 篇论文

Artificial intelligence (AI) systems will increasingly be used to cause harm as they grow more capable. In fact, AI systems are already starting to be used to automate fraudulent activities, violate human rights, create harmful fake images,…

人工智能 · 计算机科学 2023-03-30 Markus Anderljung , Julian Hazell

We propose a novel protocol for aligning artificial superintelligence (ASI) based on mutual verification among multiple isolated systems that self-modify to achieve alignment. The protocol operates by containing multiple diverse artificial…

人工智能 · 计算机科学 2025-12-01 Avraham Yair Negozio

Rapid advances in AI are beginning to reshape national security. Destabilizing AI developments could rupture the balance of power and raise the odds of great-power conflict, while widespread proliferation of capable AI hackers and…

计算机与社会 · 计算机科学 2025-04-16 Dan Hendrycks , Eric Schmidt , Alexandr Wang

If AI systems match or exceed human capabilities on a wide range of tasks, it may become difficult for humans to efficiently judge their actions -- making it hard to use human feedback to steer them towards desirable traits. One proposed…

人工智能 · 计算机科学 2025-05-26 Marie Davidsen Buhl , Jacob Pfau , Benjamin Hilton , Geoffrey Irving

The increasing use of AI technologies has led to increasing AI incidents, posing risks and causing harm to individuals, organizations, and society. This study recognizes and addresses the lack of standardized protocols for reliably and…

计算机与社会 · 计算机科学 2025-01-28 Avinash Agarwal , Manisha J Nene

Artificial intelligence (AI), although not able to currently capture the many complexities of humans, are slowly adapting to have certain capabilities of humans, many of which can revolutionize our world. AI systems, such as ChatGPT and…

计算机与社会 · 计算机科学 2024-03-26 Jay Nemec

As AI systems proliferate in society, the AI community is increasingly preoccupied with the concept of AI Safety, namely the prevention of failures due to accidents that arise from an unanticipated departure of a system's behavior from…

计算机与社会 · 计算机科学 2024-01-23 Inioluwa Deborah Raji , Roel Dobbe

We present our Balanced, Integrated and Grounded (BIG) argument for assuring the safety of AI systems. The BIG argument adopts a whole-system approach to constructing a safety case for AI systems of varying capability, autonomy and…

计算机与社会 · 计算机科学 2025-04-01 Ibrahim Habli , Richard Hawkins , Colin Paterson , Philippa Ryan , Yan Jia , Mark Sujan , John McDermid

The impact of Artificial Intelligence does not depend only on fundamental research and technological developments, but for a large part on how these systems are introduced into society and used in everyday situations. AI is changing the way…

计算机与社会 · 计算机科学 2022-05-24 Virginia Dignum

This paper provides policy recommendations to reduce extinction risks from advanced artificial intelligence (AI). First, we briefly provide background information about extinction risks from AI. Second, we argue that voluntary commitments…

人工智能 · 计算机科学 2025-09-30 Andrea Miotti

Robotics, automation, and related Artificial Intelligence (AI) systems have become pervasive bringing in concerns related to security, safety, accuracy, and trust. With growing dependency on physical robots that work in close proximity to…

密码学与安全 · 计算机科学 2023-02-17 Sudip Mittal , Jingdao Chen

Frontier AI systems are rapidly advancing in their capabilities to persuade, deceive, and influence human behaviour, with current models already demonstrating human-level persuasion and strategic deception in specific contexts. Humans are…

The advancements in generative AI inevitably raise concerns about their risks and safety implications, which, in return, catalyzes significant progress in AI safety. However, as this field continues to evolve, a critical question arises:…

计算机与社会 · 计算机科学 2025-01-14 Shanshan Han

Embedded into information systems, artificial intelligence (AI) faces security threats that exploit AI-specific vulnerabilities. This paper provides an accessible overview of adversarial attacks unique to predictive and generative AI…

密码学与安全 · 计算机科学 2025-07-01 Naoto Kiribuchi , Kengo Zenitani , Takayuki Semitsu

The promise of AI is huge. AI systems have already achieved good enough performance to be in our streets and in our homes. However, they can be brittle and unfair. For society to reap the benefits of AI systems, society needs to be able to…

人工智能 · 计算机科学 2020-02-18 Jeannette M. Wing

Artificial Intelligence (AI) is one of the disruptive technologies that is shaping the future. It has growing applications for data-driven decisions in major smart city solutions, including transportation, education, healthcare, public…

机器学习 · 计算机科学 2021-11-02 M. Humayn Kabir , Khondokar Fida Hasan , Mohammad Kamrul Hasan , Keyvan Ansari

Artificial intelligence (AI) was initially developed as an implicit moral agent to solve simple and clearly defined tasks where all options are predictable. However, it is now part of our daily life powering cell phones, cameras, watches,…

计算机与社会 · 计算机科学 2020-02-11 Mohamed Akrout , Robert Steinbauer

In the past few decades, artificial intelligence (AI) technology has experienced swift developments, changing everyone's daily life and profoundly altering the course of human society. The intention of developing AI is to benefit humans, by…

人工智能 · 计算机科学 2021-08-20 Haochen Liu , Yiqi Wang , Wenqi Fan , Xiaorui Liu , Yaxin Li , Shaili Jain , Yunhao Liu , Anil K. Jain , Jiliang Tang

Advanced AI models hold the promise of tremendous benefits for humanity, but society needs to proactively manage the accompanying risks. In this paper, we focus on what we term "frontier AI" models: highly capable foundation models that…

As a capability coming from computation, how does AI differ fundamentally from the capabilities delivered by rule-based software program? The paper examines the behavior of artificial intelligence (AI) from engineering points of view to…

人机交互 · 计算机科学 2025-11-19 Bifei Mao , Lanqing Hong