English
Related papers

Related papers: Concrete Problems in AI Safety

200 papers

How can we ensure that AI systems are aligned with human values and remain safe? We can study this problem through the frameworks of the AI assistance and the AI shutdown games. The AI assistance problem concerns designing an AI agent that…

Artificial Intelligence · Computer Science 2025-12-30 Alessio Benavoli , Alessandro Facchini , Marco Zaffalon

The complexity of dynamics in AI techniques is already approaching that of complex adaptive systems, thus curtailing the feasibility of formal controllability and reachability analysis in the context of AI safety. It follows that the…

Artificial Intelligence · Computer Science 2018-05-24 Vahid Behzadan , Arslan Munir , Roman V. Yampolskiy

As AI systems are integrated into high stakes social domains, researchers now examine how to design and operate them in a safe and ethical manner. However, the criteria for identifying and diagnosing safety risks in complex social contexts…

Computers and Society · Computer Science 2021-06-22 Roel Dobbe , Thomas Krendl Gilbert , Yonatan Mintz

Artificial intelligence (AI) is often presented as a key tool for addressing societal challenges, such as climate change. At the same time, AI's environmental footprint is expanding increasingly. This report describes the systemic…

Computers and Society · Computer Science 2025-12-16 Julian Schön , Lena Hoffmann , Nikolas Becker

This research paper explores the privacy and security threats posed to an Agentic AI system with direct access to database systems. Such access introduces significant risks, including unauthorized retrieval of sensitive information,…

Cryptography and Security · Computer Science 2024-12-10 Raihan Khan , Sayak Sarkar , Sainik Kumar Mahata , Edwin Jose

The scope of AI safety and alignment work in generative artificial intelligence (GenAI) has so far mostly been limited to harms related to: (a) discrimination and hate speech, (b) harmful/inappropriate (violent, sexual, illegal) content,…

Computers and Society · Computer Science 2026-05-06 Ilias Chalkidis , Anders Søgaard

AI scientists powered by large language models have demonstrated substantial promise in autonomously conducting experiments and facilitating scientific discoveries across various disciplines. While their capabilities are promising, these…

AI risks are typically framed around physical threats to humanity, a loss of control or an accidental error causing humanity's extinction. However, I argue in line with the gradual disempowerment thesis, that there is an underappreciated…

Computers and Society · Computer Science 2025-03-31 Joshua Krook

Artificial Intelligence began as a field probing some of the most fundamental questions of science - the nature of intelligence and the design of intelligent artifacts. But it has grown into a discipline that is deeply entwined with…

Computers and Society · Computer Science 2013-07-29 Piyush Ahuja

Integrating Artificial Intelligence (AI) into mobile and wearables offers numerous benefits at individual, societal, and environmental levels. Yet, it also spotlights concerns over emerging risks. Traditional assessments of risks and…

Human-Computer Interaction · Computer Science 2024-07-30 Marios Constantinides , Edyta Bogucka , Sanja Scepanovic , Daniele Quercia

The appreciation and utilisation of risk and uncertainty can play a key role in helping to solve some of the many ethical issues that are posed by AI. Understanding the uncertainties can allow algorithms to make better decisions by…

Computers and Society · Computer Science 2024-08-14 Nicholas Gray

With the increasing commoditization of computer vision, speech recognition and machine translation systems and the widespread deployment of learning-based back-end technologies such as digital advertising and intelligent infrastructures, AI…

This paper presents an argument that certain AI safety measures, rather than mitigating existential risk, may instead exacerbate it. Under certain key assumptions - the inevitability of AI failure, the expected correlation between an AI…

Artificial Intelligence · Computer Science 2024-06-04 Herman Cappelen , Josh Dever , John Hawthorne

The use of chatbots equipped with artificial intelligence (AI) in educational settings has increased in recent years, showing potential to support teaching and learning. However, the adoption of these technologies has raised concerns about…

Computers and Society · Computer Science 2026-03-03 Griffin Pitts , Viktoria Marcus , Sanaz Motamedi

Robot accidents are inevitable. Although rare, they have been happening since assembly-line robots were first introduced in the 1960s. But a new generation of social robots are now becoming commonplace. Often with sophisticated embedded…

Robotics · Computer Science 2020-05-18 Alan F. T. Winfield , Katie Winkle , Helena Webb , Ulrik Lyngs , Marina Jirotka , Carl Macrae

Artificial Intelligence (AI) is a double-edged sword: on one hand, AI promises to provide great advances that could benefit humanity, but on the other hand, AI poses substantial (even existential) risks. With advancements happening daily,…

Computers and Society · Computer Science 2024-02-05 Willem van der Maden , Derek Lomas , Malak Sadek , Paul Hekkert

Assuring safety of artificial intelligence (AI) applied to safety-critical systems is of paramount importance. Especially since research in the field of automated driving shows that AI is able to outperform classical approaches, to handle…

Computers and Society · Computer Science 2025-04-28 Lars Ullrich , Michael Buchholz , Klaus Dietmayer , Knut Graichen

As AI agents become more widely deployed, we are likely to see an increasing number of incidents: events involving AI agent use that directly or indirectly cause harm. For example, agents could be prompt-injected to exfiltrate private…

Computers and Society · Computer Science 2025-08-21 Carson Ezell , Xavier Roberts-Gaal , Alan Chan

Prior work has established the importance of integrating AI ethics topics into computer and data sciences curricula. We provide evidence suggesting that one of the critical objectives of AI Ethics education must be to raise awareness of AI…

Computers and Society · Computer Science 2023-10-11 Michael Feffer , Nikolas Martelaro , Hoda Heidari

This article appears as chapter 21 of Prince (2023, Understanding Deep Learning); a complete draft of the textbook is available here: http://udlbook.com. This chapter considers potential harms arising from the design and use of AI systems.…

Artificial Intelligence · Computer Science 2023-06-21 Travis LaCroix , Simon J. D. Prince