English
Related papers

Related papers: Capturing Misalignment

200 papers

We present an extensive study of the joint effects of heterogeneous social agents and their heterogeneous social links in a bounded confidence opinion dynamics model. The full phase diagram of the model is explored for two different…

Physics and Society · Physics 2023-01-24 Rémi Perrier , Hendrik Schawe , Laura Hernández

Modeling social interactions based on individual behavior has always been an area of interest, but prior literature generally presumes rational behavior. Thus, such models may miss out on capturing the effects of biases humans are…

Artificial Intelligence · Computer Science 2019-03-11 Nanda Kishore Sreenivas , Shrisha Rao

Modelling the behaviours of other agents is essential for understanding how agents interact and making effective decisions. Existing methods for agent modelling commonly assume knowledge of the local observations and chosen actions of the…

Machine Learning · Computer Science 2021-11-10 Georgios Papoudakis , Filippos Christianos , Stefano V. Albrecht

Recent advances in AI research make it increasingly plausible that artificial agents with consequential real-world impact will soon operate beyond tightly controlled environments. Ensuring that these agents are not only safe but that they…

Computers and Society · Computer Science 2025-06-10 Kevin Baum

This paper develops a dynamic equilibrium model where agents exhibit a strong form of belief heterogeneity: they disagree about zero probability events. It is shown that, somewhat surprisingly, equilibrium exists in this setting, and that…

General Finance · Quantitative Finance 2013-06-24 Martin Larsson

The value alignment problem for artificial intelligence (AI) is often framed as a purely technical or normative challenge, sometimes focused on hypothetical future systems. I argue that the problem is better understood as a structural…

Computers and Society · Computer Science 2026-04-23 Travis LaCroix

Social acceptability is an important consideration for HCI designers who develop technologies for social contexts. However, the current theoretical foundations of social acceptability research do not account for the complex interactions…

Human-Computer Interaction · Computer Science 2021-05-17 Alarith Uhde , Marc Hassenzahl

Value alignment is essential for building AI systems that can safely and reliably interact with people. However, what a person values -- and is even capable of valuing -- depends on the concepts that they are currently using to understand…

Artificial Intelligence · Computer Science 2023-11-01 Sunayana Rane , Mark Ho , Ilia Sucholutsky , Thomas L. Griffiths

Multi-agent scenarios, like Wigner's friend and Frauchiger-Renner scenarios, can show contradictory results when a non-classical formalism must deal with the knowledge between agents. Such paradoxes are described with multi-modal logic as…

Quantum Physics · Physics 2024-04-19 Sidiney B. Montanhano

Intelligent agents such as robots are increasingly deployed in real-world, safety-critical settings. It is vital that these agents are able to explain the reasoning behind their decisions to human counterparts, however, their behavior is…

Machine Learning · Computer Science 2023-09-20 Xijia Zhang , Yue Guo , Simon Stepputtis , Katia Sycara , Joseph Campbell

Auctions in which agents' payoffs are random variables have received increased attention in recent years. In particular, recent work in algorithmic mechanism design has produced mechanisms employing internal randomization, partly in…

Computer Science and Game Theory · Computer Science 2012-06-15 Shaddin Dughmi , Yuval Peres

Imitation learning, which learns agent policy by mimicking expert demonstration, has shown promising results in many applications such as medical treatment regimes and self-driving vehicles. However, it remains a difficult task to interpret…

Machine Learning · Computer Science 2024-01-31 Tianxiang Zhao , Wenchao Yu , Suhang Wang , Lu Wang , Xiang Zhang , Yuncong Chen , Yanchi Liu , Wei Cheng , Haifeng Chen

Despite investments in improving model safety, studies show that misaligned capabilities remain latent in safety-tuned models. In this work, we shed light on the mechanics of this phenomenon. First, we show that even when model generations…

Computation and Language · Computer Science 2024-08-14 Asma Ghandeharioun , Ann Yuan , Marius Guerard , Emily Reif , Michael A. Lepori , Lucas Dixon

In various economic environments, people observe other people with whom they strategically interact. We can model such information-sharing relations as an information network, and the strategic interactions as a game on the network. When…

Methodology · Statistics 2019-11-27 Nathan Canen , Jacob Schwartz , Kyungchul Song

Human behavior in interactive settings is shaped not only by individual objectives but also by shared constraints with others, such as safety. Understanding how people allocate responsibility, i.e., how much one deviates from their desired…

Multiagent Systems · Computer Science 2026-04-16 Isaac Remy , Caleb Chang , Karen Leung

We introduce a novel setting, wherein an agent needs to learn a task from a demonstration of a related task with the difference between the tasks communicated in natural language. The proposed setting allows reusing demonstrations from…

Artificial Intelligence · Computer Science 2023-01-25 Prasoon Goyal , Raymond J. Mooney , Scott Niekum

We examine the tuning of cooperative behavior in repeated multi-agent games using an analytically tractable, continuous-time, nonlinear model of opinion dynamics. Each modeled agent updates its real-valued opinion about each available…

Physics and Society · Physics 2021-11-24 Shinkyu Park , Anastasia Bizyaeva , Mari Kawakatsu , Alessio Franci , Naomi Ehrich Leonard

In dynamic settings each economic agent's choices can be revealing of her private information. This elicitation via the rationalization of observable behavior depends each agent's perception of which payoff-relevant contingencies other…

Theoretical Economics · Economics 2021-05-17 Evan Piermont , Peio Zuazo-Garin

The mental models that humans form of other agents---encapsulating human beliefs about agent goals, intentions, capabilities, and more---create an underlying basis for interaction. These mental models have the potential to affect both the…

Robotics · Computer Science 2020-01-07 Connor Brooks , Daniel Szafir

We show that in delegation problems, a principal benefits from belief misalignment vis-\`a-vis an agent when the latter can flexibly acquire costly information. The agent optimally succumbs to confirmatory learning, leading him to favor the…

Theoretical Economics · Economics 2025-07-30 Pavel Ilinov , Andrei Matveenko , Maxim Senkov , Egor Starkov