English
Related papers

Related papers: Concrete Problems in AI Safety

200 papers

Over the last decade, adversarial attack algorithms have revealed instabilities in deep learning tools. These algorithms raise issues regarding safety, reliability and interpretability in artificial intelligence; especially in high risk…

Numerical Analysis · Mathematics 2023-08-30 Desmond J. Higham

Two years after publicly launching the AI Incident Database (AIID) as a collection of harms or near harms produced by AI in the world, a backlog of "issues" that do not meet its incident ingestion criteria have accumulated in its review…

Computers and Society · Computer Science 2022-11-21 Sean McGregor , Kevin Paeth , Khoa Lam

The ability to create artificial intelligence (AI) capable of performing complex tasks is rapidly outpacing our ability to ensure the safe and assured operation of AI-enabled systems. Fortunately, a landscape of AI safety research is…

The increasing use of AI technologies has led to increasing AI incidents, posing risks and causing harm to individuals, organizations, and society. This study recognizes and addresses the lack of standardized protocols for reliably and…

Computers and Society · Computer Science 2025-01-28 Avinash Agarwal , Manisha J Nene

In the current era, people and society have grown increasingly reliant on artificial intelligence (AI) technologies. AI has the potential to drive us towards a future in which all of humanity flourishes. It also comes with substantial risks…

Computers and Society · Computer Science 2021-08-24 Lu Cheng , Kush R. Varshney , Huan Liu

Autonomous Vehicles (AV) are expected to bring considerable benefits to society, such as traffic optimization and accidents reduction. They rely heavily on advances in many Artificial Intelligence (AI) approaches and techniques. However,…

It is increasingly recognised that advances in artificial intelligence could have large and long-lasting impacts on society. However, what form those impacts will take, just how large and long-lasting they will be, and whether they will…

Computers and Society · Computer Science 2022-06-23 Sam Clarke , Jess Whittlestone

The young field of AI Safety is still in the process of identifying its challenges and limitations. In this paper, we formally describe one such impossibility result, namely Unpredictability of AI. We prove that it is impossible to…

Artificial Intelligence · Computer Science 2019-05-31 Roman V. Yampolskiy

The problem of algorithmic bias in machine learning has gained a lot of attention in recent years due to its concrete and potentially hazardous implications in society. In much the same manner, biases can also alter modern industrial and…

Machine Learning · Computer Science 2022-10-11 Laurent Risser , Agustin Picard , Lucas Hervier , Jean-Michel Loubes

As AI systems become more advanced, concerns about large-scale risks from misuse or accidents have grown. This report analyzes the technical research into safe AI development being conducted by three leading AI companies: Anthropic, Google…

Computers and Society · Computer Science 2024-09-26 Oscar Delaney , Oliver Guest , Zoe Williams

There are many goals for an AI that could become dangerous if the AI becomes superintelligent or otherwise powerful. Much work on the AI control problem has been focused on constructing AI goals that are safe even for such AIs. This paper…

Artificial Intelligence · Computer Science 2017-05-31 Stuart Armstrong , Benjamin Levinstein

As stories of human-AI interactions continue to be highlighted in the news and research platforms, the challenges are becoming more pronounced, including potential risks of overreliance, cognitive offloading, social and emotional…

Human-Computer Interaction · Computer Science 2025-10-22 Celeste Riley , Omar Al-Refai , Yadira Colunga Reyes , Eman Hammad

The exposure of security vulnerabilities in safety-aligned language models, e.g., susceptibility to adversarial attacks, has shed light on the intricate interplay between AI safety and AI security. Although the two disciplines now come…

With the introduction of Artificial Intelligence (AI) and related technologies in our daily lives, fear and anxiety about their misuse as well as the hidden biases in their creation have led to a demand for regulation to address such…

Artificial Intelligence · Computer Science 2021-04-09 The Anh Han , Tom Lenaerts , Francisco C. Santos , Luis Moniz Pereira

We sketch how developers of frontier AI systems could construct a structured rationale -- a 'safety case' -- that an AI system is unlikely to cause catastrophic outcomes through scheming. Scheming is a potential threat model where AI…

Machine learning (ML) systems are rapidly increasing in size, are acquiring new capabilities, and are increasingly deployed in high-stakes settings. As with other powerful technologies, safety for ML should be a leading research priority.…

Machine Learning · Computer Science 2022-06-20 Dan Hendrycks , Nicholas Carlini , John Schulman , Jacob Steinhardt

With ubiquitous exposure of AI systems today, we believe AI development requires crucial considerations to be deemed trustworthy. While the potential of AI systems is bountiful, though, is still unknown-as are their risks. In this work, we…

Computers and Society · Computer Science 2023-09-19 Jamell Dacon

Generative AI's humanlike qualities are driving its rapid adoption in professional domains. However, this anthropomorphic appeal raises concerns from HCI and responsible AI scholars about potential hazards and harms, such as overtrust in…

Human-Computer Interaction · Computer Science 2025-12-24 Mark Díaz , Renee Shelby , Eric Corbett , Andrew Smart

The rapid development of artificial intelligence (AI) has led to increasing concerns about the capability of AI systems to make decisions and behave responsibly. Responsible AI (RAI) refers to the development and use of AI systems that…

Software Engineering · Computer Science 2023-05-25 Boming Xia , Qinghua Lu , Harsha Perera , Liming Zhu , Zhenchang Xing , Yue Liu , Jon Whittle

Artificial intelligence (AI) advances rapidly but achieving complete human control over AI risks remains an unsolved problem, akin to driving the fast AI "train" without a "brake system." By exploring fundamental control mechanisms at key…

Computers and Society · Computer Science 2025-12-29 Yong Tao
‹ Prev 1 4 5 6 7 8 10 Next ›