English
Related papers

Related papers: Emergent Alignment via Competition

200 papers

We explore the idea of aligning an AI assistant by inverting a model of users' (unknown) preferences from observed interactions. To validate our proposal, we run proof-of-concept simulations in the economic ultimatum game, formalizing user…

Computation and Language · Computer Science 2023-12-05 Jan-Philipp Fränken , Sam Kwok , Peixuan Ye , Kanishk Gandhi , Dilip Arumugam , Jared Moore , Alex Tamkin , Tobias Gerstenberg , Noah D. Goodman

In self-consuming generative models that train on their own outputs, alignment with user preferences becomes a recursive rather than one-time process. We provide the first formal foundation for analyzing the long-term effects of such…

Machine Learning · Computer Science 2025-11-18 Ali Falahati , Mohammad Mohammadi Amiri , Kate Larson , Lukasz Golab

Fairness in language models is typically studied as a property of a single, centrally optimized model. As large language models become increasingly agentic, we propose that fairness emerges through interaction and exchange. We study this…

Computation and Language · Computer Science 2026-04-16 Sayan Kumar Chaki , Antoine Gourru , Julien Velcin

Post-training alignment optimizes language models to match human preference signals, but this objective is not equivalent to modeling observed human behavior. We compare 120 base-aligned model pairs on more than 10,000 real human decisions…

Computation and Language · Computer Science 2026-05-27 Eilam Shapira , Moshe Tennenholtz , Roi Reichart

A Nash Equilibrium (NE) is a strategy profile resilient to unilateral deviations, and is predominantly used in the analysis of multiagent systems. A downside of NE is that it is not necessarily stable against deviations by coalitions. Yet,…

Computer Science and Game Theory · Computer Science 2014-01-16 Michal Feldman , Tami Tamir

Various social contexts ranging from public goods provision to information collection can be depicted as games of strategic interactions, where a player's well-being depends on her own action as well as on the actions taken by her…

Physics and Society · Physics 2015-04-01 Giulio Cimini , Claudio Castellano , Angel Sánchez

Value alignment is essential for building AI systems that can safely and reliably interact with people. However, what a person values -- and is even capable of valuing -- depends on the concepts that they are currently using to understand…

Artificial Intelligence · Computer Science 2023-11-01 Sunayana Rane , Mark Ho , Ilia Sucholutsky , Thomas L. Griffiths

The interplay between exploration and exploitation in competitive multi-agent learning is still far from being well understood. Motivated by this, we study smooth Q-learning, a prototypical learning model that explicitly captures the…

Computer Science and Game Theory · Computer Science 2021-06-25 Stefanos Leonardos , Georgios Piliouras , Kelly Spendlove

This paper considers convex games involving multiple agents that aim to minimize their own cost functions using locally available information. A common assumption in the study of such games is that the agents are symmetric, meaning that…

Optimization and Control · Mathematics 2025-09-25 Zifan Wang , Xinlei Yi , Yi Shen , Michael M. Zavlanos , Karl H. Johansson

We study a problem where wireless service providers compete for heterogenous wireless users. The users differ in their utility functions as well as in the perceived quality of service of individual providers. We model the interaction of an…

Information Theory · Computer Science 2010-07-08 Vojislav Gajić , Jianwei Huang , Bixio Rimoldi

Reinforcement learning algorithms can train agents that solve problems in complex, interesting environments. Normally, the complexity of the trained agent is closely related to the complexity of the environment. This suggests that a highly…

Artificial Intelligence · Computer Science 2018-03-16 Trapit Bansal , Jakub Pachocki , Szymon Sidor , Ilya Sutskever , Igor Mordatch

Artificial Intelligence (AI) systems are increasingly placed in positions where their decisions have real consequences, e.g., moderating online spaces, conducting research, and advising on policy. Ensuring they operate in a safe and…

Artificial Intelligence · Computer Science 2025-05-09 Joel Z. Leibo , Alexander Sasha Vezhnevets , William A. Cunningham , Sébastien Krier , Manfred Diaz , Simon Osindero

While generic competitive systems exhibit mixtures of hierarchy and cycles, real-world systems are predominantly hierarchical. We demonstrate and extend a mechanism for hierarchy; systems with similar agents approach perfect hierarchy in…

Populations and Evolution · Quantitative Biology 2024-02-12 Christopher Cebra , Alexander Strang

The deployment of decision-making AI agents presents a critical challenge in maintaining alignment with human values or guidelines while operating in complex, dynamic environments. Agents trained solely to achieve their objectives may adopt…

Artificial Intelligence · Computer Science 2025-12-09 Dena Mujtaba , Brian Hu , Anthony Hoogs , Arslan Basharat

One obstacle to applying reinforcement learning algorithms to real-world problems is the lack of suitable reward functions. Designing such reward functions is difficult in part because the user only has an implicit understanding of the task…

Machine Learning · Computer Science 2018-11-20 Jan Leike , David Krueger , Tom Everitt , Miljan Martic , Vishal Maini , Shane Legg

Achieving convergence of multiple learning agents in general $N$-player games is imperative for the development of safe and reliable machine learning (ML) algorithms and their application to autonomous systems. Yet it is known that, outside…

Computer Science and Game Theory · Computer Science 2023-01-24 Aamal Abbas Hussain , Francesco Belardinelli , Georgios Piliouras

Algorithmic collusion has emerged as a central question in AI: Will the interaction between different AI agents deployed in markets lead to collusion? More generally, understanding how emergent behavior, be it a cartel or market dominance…

Multiagent Systems · Computer Science 2025-10-31 Ziyi Wang , Carmine Ventre , Maria Polukarov

If AI systems match or exceed human capabilities on a wide range of tasks, it may become difficult for humans to efficiently judge their actions -- making it hard to use human feedback to steer them towards desirable traits. One proposed…

Artificial Intelligence · Computer Science 2025-05-26 Marie Davidsen Buhl , Jacob Pfau , Benjamin Hilton , Geoffrey Irving

Optimizing strategic decisions (a.k.a. computing equilibrium) is key to the success of many non-cooperative multi-agent applications. However, in many real-world situations, we may face the exact opposite of this game-theoretic problem --…

Computer Science and Game Theory · Computer Science 2022-10-05 Jibang Wu , Weiran Shen , Fei Fang , Haifeng Xu

Social alignment in AI systems aims to ensure that these models behave according to established societal values. However, unlike humans, who derive consensus on value judgments through social interaction, current language models (LMs) are…

Computation and Language · Computer Science 2023-10-31 Ruibo Liu , Ruixin Yang , Chenyan Jia , Ge Zhang , Denny Zhou , Andrew M. Dai , Diyi Yang , Soroush Vosoughi