English
Related papers

Related papers: Conformity Generates Collective Misalignment in AI…

200 papers

We conduct an incentivized laboratory experiment to study people's perception of generative artificial intelligence (GenAI) alignment in the context of economic decision-making. Using a panel of economic problems spanning the domains of…

Theoretical Economics · Economics 2026-04-03 Kevin He , Ran Shorrer , Mengjia Xia

Populating our world with hyperintelligent machines obliges us to examine cognitive behaviors observed across domains that suggest autonomy may be a fundamental property of cognitive systems, and while not inherently adversarial, it…

Neurons and Cognition · Quantitative Biology 2025-06-09 Andrea Morris

As artificial intelligence (AI) becomes more powerful and widespread, the AI alignment problem - how to ensure that AI systems pursue the goals that we want them to pursue - has garnered growing attention. This article distinguishes two…

Computers and Society · Computer Science 2022-05-10 Anton Korinek , Avital Balwit

Many assumptions that underpin human concepts of identity do not hold for machine minds that can be copied, edited, or simulated. We argue that there exist many different coherent identity boundaries (e.g.\ instance, model, persona), and…

Artificial Intelligence · Computer Science 2026-03-13 Raymond Douglas , Jan Kulveit , Ondrej Havlicek , Theia Pearson-Vogel , Owen Cotton-Barratt , David Duvenaud

Large language model (LLM) agents are increasingly acting as human delegates in multi-agent environments, where a representative agent integrates diverse peer perspectives to make a final decision. Drawing inspiration from social…

Computation and Language · Computer Science 2026-05-05 Changgeon Ko , Jisu Shin , Hoyun Song , Huije Lee , Eui Jun Hwang , Jong C. Park

Beneficial societal outcomes cannot be guaranteed by aligning individual AI systems with the intentions of their operators or users. Even an AI system that is perfectly aligned to the intentions of its operating organization can lead to bad…

From school playgrounds to corporate boardrooms, status hierarchies -- rank orderings based on respect and perceived competence -- are universal features of human social organization. Language models trained on human-generated text…

Human-Computer Interaction · Computer Science 2026-01-27 Emilio Barkett

Natural and artificial collectives exhibit heterogeneities across different dimensions, contributing to the complexity of their behavior. We investigate the effect of two such heterogeneities on collective opinion dynamics: heterogeneity of…

Physics and Society · Physics 2024-02-07 Vito Mengers , Mohsen Raoufi , Oliver Brock , Heiko Hamann , Pawel Romanczuk

Large Language Models (LLMs) are increasingly instantiated as interacting agents in multi-agent systems (MAS), where collective decisions emerge through social interaction rather than independent reasoning. A fundamental yet underexplored…

Multiagent Systems · Computer Science 2026-01-12 Chen Han , Jin Tan , Bohan Yu , Wenzhen Zheng , Xijin Tang

AI alignment aims to make AI systems behave in line with human intentions and values. As AI systems grow more capable, so do risks from misalignment. To provide a comprehensive and up-to-date overview of the alignment field, in this survey,…

A key challenge for the safety of advanced AI systems is the possibility that multiple simpler agents might inadvertently form a collective agent with capabilities and goals distinct from those of any individual. More generally, determining…

Artificial Intelligence · Computer Science 2026-05-04 Frederik Hytting Jørgensen , Sebastian Weichwald , Lewis Hammond

Classification algorithms based on Artificial Intelligence (AI) are nowadays applied in high-stakes decisions in finance, healthcare, criminal justice, or education. Individuals can strategically adapt to the information gathered about…

Computer Science and Game Theory · Computer Science 2025-08-14 Marta C. Couto , Flavia Barsotti , Fernando P. Santos

Artificial Intelligence (AI) agents capable of autonomous learning and independent decision-making hold great promise for addressing complex challenges across various critical infrastructure domains, including transportation, energy…

Multiagent Systems · Computer Science 2025-07-02 Hepeng Li , Yuhong Liu , Jun Yan , Jie Gao , Xiaoou Yang

We study AI alignment through the lens of law-and-economics models of deterrence and enforcement. In these models, misconduct is not treated as an external failure, but as a strategic response to incentives: an actor weighs the gain from…

Machine Learning · Computer Science 2026-05-12 Rohit Agarwal , Joshua Lin , Mark Braverman , Elad Hazan

We investigate the spatial Public Goods Game in the presence of fitness-driven and conformity-driven agents. This framework usually considers only the former type of agents, i.e., agents that tend to imitate the strategy of their fittest…

Physics and Society · Physics 2017-08-30 Marco Alberto Javarone , Alberto Antonioni , Francesco Caravelli

The field of AI alignment aims to steer AI systems toward human goals, preferences, and ethical principles. Its contributions have been instrumental for improving the output quality, safety, and trustworthiness of today's AI models. This…

Artificial Intelligence · Computer Science 2024-11-26 Robert West , Roland Aydin

Conversational AI agents are commonly applied within single-user, turn-taking scenarios. The interaction mechanics of these scenarios are trivial: when the user enters a message, the AI agent produces a response. However, the interaction…

Large language models increasingly serve as conversational agents that adopt personas and role-play characters at user request. This capability, while valuable, raises concerns about sycophancy: the tendency to provide responses that…

Computation and Language · Computer Science 2026-04-14 Arya Shah , Deepali Mishra , Chaklam Silpasuwanchai

We investigate how peer pressure influences the opinions of Large Language Model (LLM) agents across a spectrum of cognitive commitments by embedding them in social networks where they update opinions based on peer perspectives. Our…

Computers and Society · Computer Science 2025-10-23 Aliakbar Mehdizadeh , Martin Hilbert

Traditionally, cognitive and computer scientists have viewed intelligence solipsistically, as a property of unitary agents devoid of social context. Given the success of contemporary learning algorithms, we argue that the bottleneck in…

Artificial Intelligence · Computer Science 2024-05-28 Edgar A. Duéñez-Guzmán , Suzanne Sadedin , Jane X. Wang , Kevin R. McKee , Joel Z. Leibo
‹ Prev 1 3 4 5 6 7 10 Next ›