English
Related papers

Related papers: Shutdown Safety Valves for Advanced AI

200 papers

Creating systems that are aligned with our goals is seen as a leading approach to create safe and beneficial AI in both leading AI companies and the academic field of AI safety. We defend the view that misaligned AGI - future, generally…

Computers and Society · Computer Science 2025-06-05 Max Hellrigel-Holderbaum , Leonard Dung

Artificial intelligence (AI) advances rapidly but achieving complete human control over AI risks remains an unsolved problem, akin to driving the fast AI "train" without a "brake system." By exploring fundamental control mechanisms at key…

Computers and Society · Computer Science 2025-12-29 Yong Tao

Ensuring that AI systems reliably and robustly avoid harmful or dangerous behaviours is a crucial challenge, especially for AI systems with a high degree of autonomy and general intelligence, or systems used in safety-critical contexts. In…

What does Artificial Intelligence (AI) have to contribute to health care? And what should we be looking out for if we are worried about its risks? In this paper we offer a survey, and initial evaluation, of hopes and fears about the…

Computers and Society · Computer Science 2025-05-13 Robert Sparrow , Joshua Hatherley

Artificial intelligence (AI) technologies (re-)shape modern life, driving innovation in a wide range of sectors. However, some AI systems have yielded unexpected or undesirable outcomes or have been used in questionable manners. As a…

Researchers worried about catastrophic risks from advanced AI have argued that we should expect sufficiently capable AI agents to pursue power over humanity because power is a convergent instrumental goal, something that is useful for a…

Artificial Intelligence · Computer Science 2025-06-10 Christian Tarsney

Artificial intelligence (AI) is regarded as one of the most disruptive technology of the century and with countless applications. What does it mean for radiation protection? This article describes the fundamentals of machine learning (ML)…

Machine Learning · Computer Science 2023-06-13 Sylvain Andresz , A Zéphir , Jeremy Bez , Maxime Karst , J. Danieli

United Nations set Sustainable Development Goals and this paper focuses on 7th (Affordable and Clean Energy), 9th (Industries, Innovation and Infrastructure), and 13th (Climate Action) goals. Climate change is a major concern in our…

Artificial Intelligence · Computer Science 2024-08-01 Alberto Pasqualetto , Lorenzo Serafini , Michele Sprocatti

Self-improvement is a goal currently exciting the field of AI, but is fraught with danger, and may take time to fully achieve. We advocate that a more achievable and better goal for humanity is to maximize co-improvement: collaboration…

Artificial Intelligence · Computer Science 2025-12-16 Jason Weston , Jakob Foerster

By defining the current limits (and thereby the frontiers), many boundaries are shaping, and will continue to shape, the future of Artificial Intelligence (AI). We push on these boundaries in order to make further progress into what were…

Artificial Intelligence · Computer Science 2022-05-27 Ryan Watkins , Soheil Human

This paper examines the legal implications of the explicit mentioning of automation bias (AB) in the Artificial Intelligence Act (AIA). The AIA mandates human oversight for high-risk AI systems and requires providers to enable awareness of…

Computers and Society · Computer Science 2026-02-04 Johann Laux , Hannah Ruschemeier

With edge-AI finding an increasing number of real-world applications, especially in industry, the question of functionally safe applications using AI has begun to be asked. In this body of work, we explore the issue of achieving dependable…

Machine Learning · Computer Science 2021-08-06 Hans Dermot Doran , Gianluca Ielpo , David Ganz , Michael Zapke

Defining artificial intelligence (AI) is a persistent challenge, often muddied by technical ambiguity and varying interpretations. Commonly used definitions heavily emphasize technical properties of AI but neglect the human purpose of it.…

Computers and Society · Computer Science 2024-10-21 Johannes Dahlke

Concerns around future dangers from advanced AI often centre on systems hypothesised to have intrinsic characteristics such as agent-like behaviour, strategic awareness, and long-range planning. We label this cluster of characteristics as…

Artificial Intelligence · Computer Science 2023-10-10 Kayla Matteucci , Shahar Avin , Fazl Barez , Seán Ó hÉigeartaigh

If autonomous AI systems are to be reliably safe in novel situations, they will need to incorporate general principles guiding them to recognize and avoid harmful behaviours. Such principles may need to be supported by a binding system of…

Computers and Society · Computer Science 2023-04-21 Ondrej Bajgar , Jan Horenovsky

Frontier artificial intelligence (AI) systems could pose increasing risks to public safety and security. But what level of risk is acceptable? One increasingly popular approach is to define capability thresholds, which describe AI…

Computers and Society · Computer Science 2024-06-24 Leonie Koessler , Jonas Schuett , Markus Anderljung

Modern AI systems are reaping the advantage of novel learning methods. With their increasing usage, we are realizing the limitations and shortfalls of these systems. Brittleness to minor adversarial changes in the input data, ability to…

Computers and Society · Computer Science 2020-11-05 Richa Singh , Mayank Vatsa , Nalini Ratha

Artificial Intelligence began as a field probing some of the most fundamental questions of science - the nature of intelligence and the design of intelligent artifacts. But it has grown into a discipline that is deeply entwined with…

Computers and Society · Computer Science 2013-07-29 Piyush Ahuja

Well-designed technologies that offer high levels of human control and high levels of computer automation can increase human performance, leading to wider adoption. The Human-Centered Artificial Intelligence (HCAI) framework clarifies how…

Human-Computer Interaction · Computer Science 2020-02-25 Ben Shneiderman

Artificial intelligence (AI) is increasingly of tremendous interest in the medical field. However, failures of medical AI could have serious consequences for both clinical outcomes and the patient experience. These consequences could erode…

Artificial Intelligence · Computer Science 2020-08-19 Thomas P. Quinn , Manisha Senadeera , Stephan Jacobs , Simon Coghlan , Vuong Le