English
Related papers

Related papers: The Off-Switch Game

200 papers

The dominant theories of rational choice assume logical omniscience. That is, they assume that when facing a decision problem, an agent can perform all relevant computations and determine the truth value of all relevant logical/mathematical…

Artificial Intelligence · Computer Science 2023-07-12 Caspar Oesterheld , Abram Demski , Vincent Conitzer

Consider a setting where a pre-trained agent is operating in an environment and a human operator can decide to temporarily terminate its operation and take-over for some duration of time. These kind of scenarios are common in human-machine…

Human-Computer Interaction · Computer Science 2025-05-06 Uri Menkes , Assaf Hallak , Ofra Amir

This paper focuses on a dynamic aspect of responsible autonomy, namely, to make intelligent agents be responsible at run time. That is, it considers settings where decision making by agents impinges upon the outcomes perceived by other…

Artificial Intelligence · Computer Science 2022-03-23 Munindar P. Singh

The rapid advancement of artificial intelligence (AI) systems suggests that artificial general intelligence (AGI) systems may soon arrive. Many researchers are concerned that AIs and AGIs will harm humans via intentional misuse (AI-misuse)…

Artificial Intelligence · Computer Science 2023-05-31 Catalin Mitelut , Ben Smith , Peter Vamplew

The actions of intelligent agents, such as chatbots, recommender systems, and virtual assistants are typically not fully transparent to the user. Consequently, using such an agent involves the user exposing themselves to the risk that the…

Computer Science and Game Theory · Computer Science 2020-07-23 The Anh Han , Cedric Perret , Simon T. Powers

For billions of years, evolution has been the driving force behind the development of life, including humans. Evolution endowed humans with high intelligence, which allowed us to become one of the most successful species on the planet.…

Computers and Society · Computer Science 2023-07-21 Dan Hendrycks

Many studies have shown that humans are "predictably irrational": they do not act in a fully rational way, but their deviations from rational behavior are quite systematic. Our goal is to see the extent to which we can explain and justify…

Computer Science and Game Theory · Computer Science 2023-07-27 Xinming Liu , Joseph Y. Halpern

Online information ecosystems are now central to our everyday social interactions. Of the many opportunities and challenges this presents, the capacity for artificial agents to shape individual and collective human decision-making in such…

Populations and Evolution · Quantitative Biology 2023-12-07 Theodor Cimpeanu , Alexander J. Stewart

Reinforcement learning agents have been mostly developed and evaluated under the assumption that they will operate in a fully autonomous manner -- they will take all actions. In this work, our goal is to develop algorithms that, by learning…

Machine Learning · Computer Science 2023-07-04 Vahid Balazadeh , Abir De , Adish Singla , Manuel Gomez-Rodriguez

We study partially observable assistance games (POAGs), a model of the human-AI value alignment problem which allows the human and the AI assistant to have partial observations. Motivated by concerns of AI deception, we study a…

Artificial Intelligence · Computer Science 2025-08-12 Scott Emmons , Caspar Oesterheld , Vincent Conitzer , Stuart Russell

Reinforcement learning (RL) agents optimize only the features specified in a reward function and are indifferent to anything left out inadvertently. This means that we must not only specify what to do, but also the much larger space of what…

Machine Learning · Computer Science 2019-04-22 Rohin Shah , Dmitrii Krasheninnikov , Jordan Alexander , Pieter Abbeel , Anca Dragan

Natural Immune system plays a vital role in the survival of the all living being. It provides a mechanism to defend itself from external predates making it consistent systems, capable of adapting itself for survival incase of changes. The…

Artificial Intelligence · Computer Science 2011-03-11 Tejbanta Singh Chingtham , G. Sahoo , M. K. Ghose

If AI systems match or exceed human capabilities on a wide range of tasks, it may become difficult for humans to efficiently judge their actions -- making it hard to use human feedback to steer them towards desirable traits. One proposed…

Artificial Intelligence · Computer Science 2025-05-26 Marie Davidsen Buhl , Jacob Pfau , Benjamin Hilton , Geoffrey Irving

Auctions in which agents' payoffs are random variables have received increased attention in recent years. In particular, recent work in algorithmic mechanism design has produced mechanisms employing internal randomization, partly in…

Computer Science and Game Theory · Computer Science 2012-06-15 Shaddin Dughmi , Yuval Peres

Recent developments in artificial intelligence (AI) have permeated through an array of different immersive environments, including virtual, augmented, and mixed realities. AI brings a wealth of potential that centers on its ability to…

Human-Computer Interaction · Computer Science 2024-05-10 Wangfan Li , Rohit Mallick , Carlos Toxtli-Hernandez , Christopher Flathmann , Nathan J. McNeese

The theory of rational choice assumes that when people make decisions they do so in order to maximize their utility. In order to achieve this goal they ought to use all the information available and consider all the choices available to…

Artificial Intelligence · Computer Science 2017-04-07 Tshilidzi Marwala

The rational choice theory is based on this idea that people rationally pursue goals for increasing their personal interests. In most conditions, the behavior of an actor is not independent of the person and others' behavior. Here, we…

Econometrics · Economics 2018-02-27 Madjid Eshaghi Gordji , Gholamreza Askari

Corrigibility of autonomous agents is an under explored part of system design, with previous work focusing on single agent systems. It has been suggested that uncertainty over the human preferences acts to keep the agents corrigible, even…

Computer Science and Game Theory · Computer Science 2025-01-10 Edmund Dable-Heath , Boyko Vodenicharski , James Bishop

For AI systems to be useful to humans, they must understand and act in accordance with our values and preferences. Since specifying preferences is a hard task, inverse reinforcement learning (IRL) aims to develop methods that allow for…

Artificial Intelligence · Computer Science 2026-05-12 Karim Abdel Sadek , Mark Bedaywi , Rhys Gould , Stuart Russell

We investigate whether and why people might adjust compensation for workers who use AI tools. Across 13 studies (N = 4,956), participants consistently lowered compensation for workers who used AI compared to those who did not. This "AI…

General Economics · Economics 2026-03-06 Jin Kim , Shane Schweitzer , David De Cremer , Christoph Riedl