中文
相关论文

相关论文: Shutdown Safety Valves for Advanced AI

200 篇论文

Creating systems that are aligned with our goals is seen as a leading approach to create safe and beneficial AI in both leading AI companies and the academic field of AI safety. We defend the view that misaligned AGI - future, generally…

计算机与社会 · 计算机科学 2025-06-05 Max Hellrigel-Holderbaum , Leonard Dung

Artificial intelligence (AI) advances rapidly but achieving complete human control over AI risks remains an unsolved problem, akin to driving the fast AI "train" without a "brake system." By exploring fundamental control mechanisms at key…

计算机与社会 · 计算机科学 2025-12-29 Yong Tao

Ensuring that AI systems reliably and robustly avoid harmful or dangerous behaviours is a crucial challenge, especially for AI systems with a high degree of autonomy and general intelligence, or systems used in safety-critical contexts. In…

What does Artificial Intelligence (AI) have to contribute to health care? And what should we be looking out for if we are worried about its risks? In this paper we offer a survey, and initial evaluation, of hopes and fears about the…

计算机与社会 · 计算机科学 2025-05-13 Robert Sparrow , Joshua Hatherley

Artificial intelligence (AI) technologies (re-)shape modern life, driving innovation in a wide range of sectors. However, some AI systems have yielded unexpected or undesirable outcomes or have been used in questionable manners. As a…

Researchers worried about catastrophic risks from advanced AI have argued that we should expect sufficiently capable AI agents to pursue power over humanity because power is a convergent instrumental goal, something that is useful for a…

人工智能 · 计算机科学 2025-06-10 Christian Tarsney

Artificial intelligence (AI) is regarded as one of the most disruptive technology of the century and with countless applications. What does it mean for radiation protection? This article describes the fundamentals of machine learning (ML)…

机器学习 · 计算机科学 2023-06-13 Sylvain Andresz , A Zéphir , Jeremy Bez , Maxime Karst , J. Danieli

United Nations set Sustainable Development Goals and this paper focuses on 7th (Affordable and Clean Energy), 9th (Industries, Innovation and Infrastructure), and 13th (Climate Action) goals. Climate change is a major concern in our…

人工智能 · 计算机科学 2024-08-01 Alberto Pasqualetto , Lorenzo Serafini , Michele Sprocatti

Self-improvement is a goal currently exciting the field of AI, but is fraught with danger, and may take time to fully achieve. We advocate that a more achievable and better goal for humanity is to maximize co-improvement: collaboration…

人工智能 · 计算机科学 2025-12-16 Jason Weston , Jakob Foerster

By defining the current limits (and thereby the frontiers), many boundaries are shaping, and will continue to shape, the future of Artificial Intelligence (AI). We push on these boundaries in order to make further progress into what were…

人工智能 · 计算机科学 2022-05-27 Ryan Watkins , Soheil Human

This paper examines the legal implications of the explicit mentioning of automation bias (AB) in the Artificial Intelligence Act (AIA). The AIA mandates human oversight for high-risk AI systems and requires providers to enable awareness of…

计算机与社会 · 计算机科学 2026-02-04 Johann Laux , Hannah Ruschemeier

With edge-AI finding an increasing number of real-world applications, especially in industry, the question of functionally safe applications using AI has begun to be asked. In this body of work, we explore the issue of achieving dependable…

机器学习 · 计算机科学 2021-08-06 Hans Dermot Doran , Gianluca Ielpo , David Ganz , Michael Zapke

Defining artificial intelligence (AI) is a persistent challenge, often muddied by technical ambiguity and varying interpretations. Commonly used definitions heavily emphasize technical properties of AI but neglect the human purpose of it.…

计算机与社会 · 计算机科学 2024-10-21 Johannes Dahlke

Concerns around future dangers from advanced AI often centre on systems hypothesised to have intrinsic characteristics such as agent-like behaviour, strategic awareness, and long-range planning. We label this cluster of characteristics as…

人工智能 · 计算机科学 2023-10-10 Kayla Matteucci , Shahar Avin , Fazl Barez , Seán Ó hÉigeartaigh

If autonomous AI systems are to be reliably safe in novel situations, they will need to incorporate general principles guiding them to recognize and avoid harmful behaviours. Such principles may need to be supported by a binding system of…

计算机与社会 · 计算机科学 2023-04-21 Ondrej Bajgar , Jan Horenovsky

Frontier artificial intelligence (AI) systems could pose increasing risks to public safety and security. But what level of risk is acceptable? One increasingly popular approach is to define capability thresholds, which describe AI…

计算机与社会 · 计算机科学 2024-06-24 Leonie Koessler , Jonas Schuett , Markus Anderljung

Modern AI systems are reaping the advantage of novel learning methods. With their increasing usage, we are realizing the limitations and shortfalls of these systems. Brittleness to minor adversarial changes in the input data, ability to…

计算机与社会 · 计算机科学 2020-11-05 Richa Singh , Mayank Vatsa , Nalini Ratha

Artificial Intelligence began as a field probing some of the most fundamental questions of science - the nature of intelligence and the design of intelligent artifacts. But it has grown into a discipline that is deeply entwined with…

计算机与社会 · 计算机科学 2013-07-29 Piyush Ahuja

Well-designed technologies that offer high levels of human control and high levels of computer automation can increase human performance, leading to wider adoption. The Human-Centered Artificial Intelligence (HCAI) framework clarifies how…

人机交互 · 计算机科学 2020-02-25 Ben Shneiderman

Artificial intelligence (AI) is increasingly of tremendous interest in the medical field. However, failures of medical AI could have serious consequences for both clinical outcomes and the patient experience. These consequences could erode…

人工智能 · 计算机科学 2020-08-19 Thomas P. Quinn , Manisha Senadeera , Stephan Jacobs , Simon Coghlan , Vuong Le