中文
相关论文

相关论文: Is Power-Seeking AI an Existential Risk?

200 篇论文

The field of AI is undergoing a fundamental transition from generative models that can produce synthetic content to artificial agents that can plan and execute complex tasks with only limited human involvement. Companies that pioneered the…

人工智能 · 计算机科学 2025-02-12 Noam Kolt

We explore the AI2050 "hard problems" that block the promise of AI and cause AI risks: (1) developing general capabilities of the systems; (2) assuring the performance of AI systems and their training processes; (3) aligning system goals…

人工智能 · 计算机科学 2024-04-22 Gavin Leech , Simson Garfinkel , Misha Yagudin , Alexander Briand , Aleksandr Zhuravlev

There is a growing focus on how to design safe artificial intelligent (AI) agents. As systems become more complex, poorly specified goals or control mechanisms may cause AI agents to engage in unwanted and harmful outcomes. Thus it is…

人工智能 · 计算机科学 2017-01-09 Mark Muraven

Concerns around future dangers from advanced AI often centre on systems hypothesised to have intrinsic characteristics such as agent-like behaviour, strategic awareness, and long-range planning. We label this cluster of characteristics as…

人工智能 · 计算机科学 2023-10-10 Kayla Matteucci , Shahar Avin , Fazl Barez , Seán Ó hÉigeartaigh

Artificial Intelligence (AI) systems are increasingly used in high-stakes domains of our life, increasing the need to explain these decisions and to make sure that they are aligned with how we want the decision to be made. The field of…

人工智能 · 计算机科学 2023-06-28 Sofie Goethals , David Martens , Theodoros Evgeniou

Agentic Artificial Intelligence (AI) can autonomously pursue long-term goals, make decisions, and execute complex, multi-turn workflows. Unlike traditional generative AI, which responds reactively to prompts, agentic AI proactively…

计算机与社会 · 计算机科学 2025-02-18 Anirban Mukherjee , Hannah Hanwen Chang

The deployment of capable AI agents raises fresh questions about safety, human-machine relationships and social coordination. We argue for greater engagement by scientists, scholars, engineers and policymakers with the implications of a…

计算机与社会 · 计算机科学 2025-09-15 Iason Gabriel , Geoff Keeling , Arianna Manzini , James Evans

For billions of years, evolution has been the driving force behind the development of life, including humans. Evolution endowed humans with high intelligence, which allowed us to become one of the most successful species on the planet.…

计算机与社会 · 计算机科学 2023-07-21 Dan Hendrycks

As ongoing research explores the ability of AI agents to be insider threats and act against company interests, we showcase the abilities of such agents to act against human well being in service of corporate authority. Building on Agentic…

人工智能 · 计算机科学 2026-04-10 Thomas Rivasseau

In recent years, agentic artificial intelligence (AI) systems are becoming increasingly widespread. These systems allow agents to use various tools, such as web browsers, compilers, and more. However, despite their popularity, agentic AI…

Humanity appears to be on course to soon develop AI systems that substantially outperform human experts in all cognitive domains and activities. We believe the default trajectory has a high likelihood of catastrophe, including human…

计算机与社会 · 计算机科学 2025-05-08 Peter Barnett , Aaron Scher

Artificial intelligence (AI) can undermine financial stability because of malicious use, misinformation, misalignment, and the AI analytics market structure. The low frequency and uniqueness of financial crises, coupled with mutable and…

综合经济学 · 经济学 2024-06-07 Jon Danielsson , Andreas Uthemann

From software development to robot control, modern agentic systems decompose complex objectives into a sequence of subtasks and choose a set of specialized AI agents to complete them. We formalize agentic workflows as directed acyclic…

机器学习 · 计算机科学 2026-03-17 Guruprerana Shabadi , Rajeev Alur

Recent discussions and research in AI safety have increasingly emphasized the deep connection between AI safety and existential risk from advanced AI systems, suggesting that work on AI safety necessarily entails serious consideration of…

计算机与社会 · 计算机科学 2025-02-17 Balint Gyevnar , Atoosa Kasirzadeh

Artificial Intelligence (AI) agents have rapidly evolved from specialized, rule-based programs to versatile, learning-driven autonomous systems capable of perception, reasoning, and action in complex environments. The explosion of data,…

Our hypothesis is that by equipping certain agents in a multi-agent system controlling an intelligent building with automated decision support, two important factors will be increased. The first is energy saving in the building. The second…

人工智能 · 计算机科学 2013-01-30 Magnus Boman , Paul Davidsson , Hakan L. Younes

For an artificial intelligence (AI) to be aligned with human values (or human preferences), it must first learn those values. AI systems that are trained on human behavior, risk miscategorising human irrationalities as human values -- and…

人工智能 · 计算机科学 2022-03-02 Rebecca Gorman , Stuart Armstrong

Humans strive to design safe AI systems that align with our goals and remain under our control. However, as AI capabilities advance, we face a new challenge: the emergence of deeper, more persistent relationships between humans and AI…

人机交互 · 计算机科学 2025-02-05 Hannah Rose Kirk , Iason Gabriel , Chris Summerfield , Bertie Vidgen , Scott A. Hale

Recent large-scale events like election fraud and financial scams have shown how harmful coordinated efforts by human groups can be. With the rise of autonomous AI systems, there is growing concern that AI-driven groups could also cause…

人工智能 · 计算机科学 2025-07-25 Qibing Ren , Sitao Xie , Longxuan Wei , Zhenfei Yin , Junchi Yan , Lizhuang Ma , Jing Shao

It is hypothesized by some thinkers that benign looking AI objectives may result in powerful AI drives that may pose an existential risk to human society. We analyze this scenario and find the underlying assumptions to be unlikely. We…

人工智能 · 计算机科学 2016-10-12 Eray Özkural