English
Related papers

Related papers: Capturing Misalignment

200 papers

When developing AI systems that interact with humans, it is essential to design both a system that can understand humans, and a system that humans can understand. Most deep network based agent-modeling approaches are 1) not interpretable…

Machine Learning · Computer Science 2021-07-14 Ini Oguntola , Dana Hughes , Katia Sycara

To coordinate with other agents in its environment, an agent needs models of what the other agents are trying to do. When communication is impossible or expensive, this information must be acquired indirectly via plan recognition. Typical…

Artificial Intelligence · Computer Science 2013-02-28 Marcus J. Huber , Edmund H. Durfee , Michael P. Wellman

Based on criteria of mathematical simplicity and consistency with empirical market data, a model with volatility driven by fractional noise has been constructed which provides a fairly accurate mathematical parametrization of the data.…

Statistical Finance · Quantitative Finance 2010-08-31 R. Vilela Mendes

Consider an imitation learning problem that the imitator and the expert have different dynamics models. Most of the current imitation learning methods fail because they focus on imitating actions. We propose a novel state alignment-based…

Machine Learning · Computer Science 2019-11-26 Fangchen Liu , Zhan Ling , Tongzhou Mu , Hao Su

Active inference is a first principles approach for understanding the brain in particular, and sentient agents in general, with the single imperative of minimizing free energy. As such, it provides a computational account for modelling…

Computer Vision and Pattern Recognition · Computer Science 2022-09-20 Stefano Ferraro , Toon Van de Maele , Pietro Mazzaglia , Tim Verbelen , Bart Dhoedt

Simultaneous reproduction of all financial stylized facts is so difficult that most existing stochastic process-based and agent-based models are unable to achieve the goal. In this study, by extending the decision-making structure of…

Statistical Finance · Quantitative Finance 2019-05-22 Kei Katahira , Yu Chen , Gaku Hashimoto , Hiroshi Okuda

This paper explores the mechanistic interpretability of reinforcement learning (RL) agents through an analysis of a neural network trained on procedural maze environments. By dissecting the network's inner workings, we identified…

Machine Learning · Computer Science 2024-11-05 Tristan Trim , Triston Grayston

This paper develops a new approach for estimating an interpretable, relational model of a black-box autonomous agent that can plan and act. Our main contributions are a new paradigm for estimating such models using a minimal query interface…

Artificial Intelligence · Computer Science 2021-04-12 Pulkit Verma , Shashank Rao Marpally , Siddharth Srivastava

In multiagent systems autonomous agents interact with each other to achieve individual and collective goals. Typical interactions concern negotiation and agreement on resource exchanges. Modeling and formalizing these agreements pose…

Logic in Computer Science · Computer Science 2024-08-20 Lorenzo Ceragioli , Pierpaolo Degano , Letterio Galletta , Luca Viganò

Following a long tradition of physicists who have noticed that the Ising model provides a general background to build realistic models of social interactions, we study a model of financial price dynamics resulting from the collective…

Statistical Mechanics · Physics 2008-12-02 Didier Sornette , Wei-Xing Zhou

We investigate opinion formation in a kinetic exchange opinion model, where opinions are represented by numbers in the real interval $[-1,1]$ and agents are typified by the individual degree of conviction about the opinion that they…

Physics and Society · Physics 2016-02-25 Allan R. Vieira , Celia Anteneodo , Nuno Crokidakis

Common knowledge/belief in rationality is the traditional standard assumption in analysing interaction among agents. This paper proposes a graph-based language for capturing significantly more complicated structures of higher-order beliefs…

Artificial Intelligence · Computer Science 2024-12-13 Qi Shi , Pavel Naumov

We examine the long-term behavior of a Bayesian agent who has a misspecified belief about the time lag between actions and feedback, and learns about the payoff consequences of his actions over time. Misspecified beliefs about time lags…

Theoretical Economics · Economics 2020-12-15 Yingkai Li , Harry Pei

Game-theoretic interactions with AI agents could differ from traditional human-human interactions in various ways. One such difference is that it may be possible to simulate an AI agent (for example because its source code is known), which…

Computer Science and Game Theory · Computer Science 2024-03-21 Vojtech Kovarik , Caspar Oesterheld , Vincent Conitzer

As the complexity of AI systems and their interactions with the world increases, generating explanations for their behaviour is important for safely deploying AI. For agents, the most natural abstractions for predicting behaviour attribute…

Artificial Intelligence · Computer Science 2025-06-05 Alexis Bellot , Jonathan Richens , Tom Everitt

This paper studies when strategic understanding acquired in one mechanism can be transferred to another. We introduce a framework in which agents' knowledge is represented as a set of payoff comparisons they can make, and use it to…

Theoretical Economics · Economics 2026-05-14 Joseph Feffer , Filip Tokarski

An agent chooses an action based on her private information and a recommendation from an informed but potentially misaligned adviser. With a known probability, the adviser truthfully reports his signal; with the remaining probability, he…

Theoretical Economics · Economics 2026-03-20 Piotr Dworczak , Alex Smolin

Familiarity with a simulation platform can seduce modellers into accepting untested assumptions for convenience of implementation. These assumptions may have consequences greater than commonly suspected, and it is important that modellers…

Quantitative Methods · Quantitative Biology 2014-02-28 Jerome K Vanclay

The AI alignment problem, which focusses on ensuring that artificial intelligence (AI), including AGI and ASI, systems act according to human values, presents profound challenges. With the progression from narrow AI to Artificial General…

Artificial Intelligence · Computer Science 2025-07-25 Alberto Hernández-Espinosa , Felipe S. Abrahão , Olaf Witkowski , Hector Zenil

From marketing to politics, exploitation of incomplete information through selective communication of arguments is ubiquitous. In this work, we focus on development of an argumentation-theoretic model for manipulable multi-agent…

Artificial Intelligence · Computer Science 2019-09-17 Ryuta Arisaka , Makoto Hagiwara , Takayuki Ito