English
Related papers

Related papers: Consequences of Misaligned AI

200 papers

Principal-agent problems arise when one party acts on behalf of another, leading to conflicts of interest. The economic literature has extensively studied principal-agent problems, and recent work has extended this to more complex scenarios…

Artificial Intelligence · Computer Science 2024-01-02 Omer Ben-Porat , Yishay Mansour , Michal Moshkovitz , Boaz Taitler

The optimal objective is a fundamental aspect of reinforcement learning (RL), as it determines how policies are evaluated and optimized. While total return maximization is the ideal objective in RL, discounted return maximization is the…

Machine Learning · Computer Science 2025-03-19 Shuyu Yin , Fei Wen , Peilin Liu , Tao Luo

A principal who values an object allocates it to one or more agents. Agents learn private information (signals) from an information designer about the allocation payoff to the principal. Monetary transfer is not available but the principal…

Theoretical Economics · Economics 2022-10-31 Yi-Chun Chen , Gaoji Hu , Xiangqian Yang

Incentives are more likely to elicit desired outcomes when they are designed based on accurate models of agents' strategic behavior. A growing literature, however, suggests that people do not quite behave like standard economic agents in a…

Computer Science and Game Theory · Computer Science 2014-06-09 Arpita Ghosh , Robert Kleinberg

When should we delegate decisions to AI systems? While the value alignment literature has developed techniques for shaping AI values, less attention has been paid to how to determine, under uncertainty, when imperfect alignment is good…

Artificial Intelligence · Computer Science 2025-12-23 Daniel A. Herrmann , Abinav Chari , Isabelle Qian , Sree Sharvesh , B. A. Levinstein

We consider the problem of incentivising desirable behaviours in multi-agent systems by way of taxation schemes. Our study employs the concurrent games model: in this model, each agent is primarily motivated to seek the satisfaction of a…

Computer Science and Game Theory · Computer Science 2023-07-12 David Hyland , Julian Gutierrez , Michael Wooldridge

AI agents are commonly trained with large datasets of demonstrations of human behavior. However, not all behaviors are equally safe or desirable. Desired characteristics for an AI agent can be expressed by assigning desirability scores,…

Machine Learning · Computer Science 2024-05-08 Tim Franzmeyer , Edith Elkind , Philip Torr , Jakob Foerster , Joao Henriques

We study the problem of cooperative multi-agent reinforcement learning with a single joint reward signal. This class of learning problems is difficult because of the often large combined action and observation spaces. In the fully…

Consider a principal who wants to search through a space of stochastic solutions for one maximizing their utility. If the principal cannot conduct this search on their own, they may instead delegate this problem to an agent with distinct…

Computer Science and Game Theory · Computer Science 2024-11-04 Curtis Bechtel , Shaddin Dughmi

Ensuring that AI agents behave safely and beneficially when interacting with other parties has emerged as one of the central challenges of modern AI safety. While mechanism design, as the theory of designing rules to align individual and…

Computer Science and Game Theory · Computer Science 2026-05-12 Xuanqiang Angelo Huang , Charlie Tharas , Samuele Marro , Van Q. Truong , Bernhard Schölkopf , Emanuele La Malfa , Zhijing Jin

We propose a multi-agent distributed reinforcement learning algorithm that balances between potentially conflicting short-term reward and sparse, delayed long-term reward, and learns with partial information in a dynamic environment. We…

Machine Learning · Computer Science 2022-04-06 Jing Tan , Ramin Khalili , Holger Karl

Continually solving new, unsolved tasks is the key to learning diverse behaviors. Through reinforcement learning (RL), we have made massive strides towards solving tasks that have a single goal. However, in the multi-task domain, where an…

Machine Learning · Computer Science 2020-06-18 Yunzhi Zhang , Pieter Abbeel , Lerrel Pinto

In this paper, we consider the problem of a Principal aiming at designing a reward function for a population of heterogeneous agents. We construct an incentive based on the ranking of the agents, so that a competition among the latter is…

Optimization and Control · Mathematics 2026-04-28 Clémence Alasseur , Erhan Bayraktar , Roxana Dumitrescu , Quentin Jacquet

Motivated by a number of real-world applications from domains like healthcare and sustainable transportation, in this paper we study a scenario of repeated principal-agent games within a multi-armed bandit (MAB) framework, where: the…

Machine Learning · Computer Science 2023-05-09 Ilgin Dogan , Zuo-Jun Max Shen , Anil Aswani

A hallmark property of explainable AI models is the ability to teach other agents, communicating knowledge of how to perform a task. While Large Language Models perform complex reasoning by generating explanations for their predictions, it…

Computation and Language · Computer Science 2023-11-15 Swarnadeep Saha , Peter Hase , Mohit Bansal

We consider the classic principal-agent model of contract theory, in which a principal designs an outcome-dependent compensation scheme to incentivize an agent to take a costly and unobservable action. When all of the model…

Computer Science and Game Theory · Computer Science 2020-08-11 Paul Dütting , Tim Roughgarden , Inbal Talgam-Cohen

The AI-alignment problem arises when there is a discrepancy between the goals that a human designer specifies to an AI learner and a potential catastrophic outcome that does not reflect what the human designer really wants. We argue that a…

Machine Learning · Computer Science 2020-04-10 Shai Shalev-Shwartz , Shaked Shammah , Amnon Shashua

Overcoming the impact of selfish behavior of rational players in multiagent systems is a fundamental problem in game theory. Without any intervention from a central agent, strategic users take actions in order to maximize their personal…

Computer Science and Game Theory · Computer Science 2024-09-06 Maria-Florina Balcan , Matteo Pozzi , Dravyansh Sharma

AI systems increasingly assist human decision making by producing preliminary assessments of complex inputs. However, such AI-generated assessments can often be noisy or systematically biased, raising a central question: how should costly…

Machine Learning · Statistics 2026-03-17 Lezhi Tan , Naomi Sagan , Lihua Lei , Jose Blanchet

In a multi-party machine learning system, different parties cooperate on optimizing towards better models by sharing data in a privacy-preserving way. A major challenge in learning is the incentive issue. For example, if there is…

Multiagent Systems · Computer Science 2020-08-11 Mengjing Chen , Yang Liu , Weiran Shen , Yiheng Shen , Pingzhong Tang , Qiang Yang