English
Related papers

Related papers: Consequentialist Objectives and Catastrophe

200 papers

Human beings are able to master a variety of knowledge and skills with ongoing learning. By contrast, dramatic performance degradation is observed when new tasks are added to an existing neural network model. This phenomenon, termed as…

Machine Learning · Computer Science 2019-10-25 Xin Yao , Tianchi Huang , Chenglei Wu , Rui-Xiao Zhang , Lifeng Sun

The trustworthiness of AI decision-making systems is increasingly important. A key feature of such systems is the ability to provide recommendations for how an individual may reverse a negative decision, a problem known as algorithmic…

Artificial Intelligence · Computer Science 2026-05-13 Drago Plecko , Collin Wang , Elias Bareinboim

Reinforcement learning algorithms are generally designed to maximize the expected return across a population. However, a policy that is optimal on average may be suboptimal for certain individuals, leading to potential safety concerns. To…

Machine Learning · Statistics 2026-05-26 Jingyi Li , Peng Wu , Chengchun Shi

The rise of AI has transformed the software and hardware landscape, enabling powerful capabilities through specialized infrastructures, large-scale data storage, and advanced hardware. However, these innovations introduce unique attack…

Cryptography and Security · Computer Science 2025-08-29 Michael R Smith , Joe Ingram

We identify a distinct motive for search, termed catalytic exploration, where agents rationally explore alternatives they expect to reject to resolve uncertainty about the status quo. By decomposing option value into switching and catalytic…

Theoretical Economics · Economics 2025-11-25 Zeyu He

We systematically evaluate the quality of widely used adversarial safety datasets from two perspectives: in isolation and in practice. In isolation, we examine how well these datasets reflect real-world adversarial attacks based on three…

Cryptography and Security · Computer Science 2026-04-24 Shahriar Golchin , Marc Wetter

The appreciation and utilisation of risk and uncertainty can play a key role in helping to solve some of the many ethical issues that are posed by AI. Understanding the uncertainties can allow algorithms to make better decisions by…

Computers and Society · Computer Science 2024-08-14 Nicholas Gray

In many contexts, lying -- the use of verbal falsehoods to deceive -- is harmful. While lying has traditionally been a human affair, AI systems that make sophisticated verbal statements are becoming increasingly prevalent. This raises the…

Computers and Society · Computer Science 2021-10-14 Owain Evans , Owen Cotton-Barratt , Lukas Finnveden , Adam Bales , Avital Balwit , Peter Wills , Luca Righetti , William Saunders

There is a substantial and ever-growing corpus of evidence and literature exploring the impacts of Artificial intelligence (AI) technologies on society, politics, and humanity as a whole. A separate, parallel body of work has explored…

Computers and Society · Computer Science 2022-09-23 Benjamin S. Bucknall , Shiri Dori-Hacohen

The rapid development of AI systems poses unprecedented risks, including loss of control, misuse, geopolitical instability, and concentration of power. To navigate these risks and avoid worst-case outcomes, governments may proactively…

Artificial Intelligence · Computer Science 2025-07-15 Peter Barnett , Aaron Scher , David Abecassis

When there exists uncertainty, AI machines are designed to make decisions so as to reach the best expected outcomes. Expectations are based on true facts about the objective environment the machines interact with, and those facts can be…

Machine Learning · Computer Science 2024-07-09 Jinsook Kim

AI-driven outcomes can be challenging for end-users to understand. Explanations can address two key questions: "Why this outcome?" (factual) and "Why not another?" (counterfactual). While substantial efforts have been made to formalize…

Artificial Intelligence · Computer Science 2025-03-21 Suryani Lim , Henri Prade , Gilles Richard

The ability to learn tasks in a sequential fashion is crucial to the development of artificial intelligence. Neural networks are not, in general, capable of this and it has been widely thought that catastrophic forgetting is an inevitable…

Algorithmic recourse provides explanations that help users overturn an unfavorable decision by a machine learning system. But so far very little attention has been paid to whether providing recourse is beneficial or not. We introduce an…

Machine Learning · Computer Science 2024-03-04 Hidde Fokkema , Damien Garreau , Tim van Erven

Supervised Continual learning involves updating a deep neural network (DNN) from an ever-growing stream of labeled data. While most work has focused on overcoming catastrophic forgetting, one of the major motivations behind continual…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Md Yousuf Harun , Jhair Gallardo , Tyler L. Hayes , Christopher Kanan

As the prevalence and everyday use of machine learning algorithms, along with our reliance on these algorithms grow dramatically, so do the efforts to attack and undermine these algorithms with malicious intent, resulting in a growing…

Machine Learning · Statistics 2018-02-22 Christopher Frederickson , Michael Moore , Glenn Dawson , Robi Polikar

As a result of rapidly accelerating AI capabilities, over the past year, national governments and multinational bodies have announced efforts to address safety, security and ethics issues related to AI models. One high priority among these…

Computers and Society · Computer Science 2026-05-19 Jaspreet Pannu , Doni Bloomfield , Alex Zhu , Robert MacKnight , Gabe Gomes , Anita Cicero , Thomas V. Inglesby

Recent empirical results have demonstrated that training large language models (LLMs) with negative-only feedback can match or exceed standard reinforcement learning from human feedback (RLHF). Negative Sample Reinforcement achieves parity…

Artificial Intelligence · Computer Science 2026-03-18 Quan Cheng

The leading AI companies are increasingly focused on building generalist AI agents -- systems that can autonomously plan, act, and pursue goals across almost all tasks that humans can perform. Despite how useful these systems might be,…

Can AI agents deal with hard choices -- cases where options are incommensurable because multiple objectives are pursued simultaneously? Adopting a technologically engaged approach distinct from existing philosophical literature, I submit…

Artificial Intelligence · Computer Science 2026-04-21 Kangyu Wang