English
Related papers

Related papers: The Limits of Predicting Agents from Behaviour

200 papers

As AI systems are increasingly incorporated into domains where human behavior has set the norm, a challenge for AI governance and AI alignment research is to regulate their behavior in a way that is useful and constructive for society. One…

Computers and Society · Computer Science 2024-06-10 Sunayana Rane

Could belief in AI predictions be just another form of superstition? This study investigates psychological factors that influence belief in AI predictions about personal behavior, comparing it to belief in astrology- and personality-based…

Human-Computer Interaction · Computer Science 2024-12-20 Eunhae Lee , Pat Pataranutaporn , Judith Amores , Pattie Maes

Artificial intelligence (AI) systems powered by large language models have become increasingly prevalent in modern society, enabling a wide range of applications through natural language interaction. As AI agents proliferate in our daily…

Machine Learning · Computer Science 2025-03-24 J. M. Diederik Kruijssen , Nicholas Emmons

Power-seeking behavior is a key source of risk from advanced AI, but our theoretical understanding of this phenomenon is relatively limited. Building on existing theoretical results demonstrating power-seeking incentives for most reward…

Artificial Intelligence · Computer Science 2023-04-14 Victoria Krakovna , Janos Kramar

The concept of the 'agent' has profoundly shaped Artificial Intelligence (AI) research, guiding development from foundational theories to contemporary applications like Large Language Model (LLM)-based systems. This paper critically…

Artificial Intelligence · Computer Science 2025-09-16 Jesse Gardner , Vladimir A. Baulin

The aim of my Ph.D. thesis concerns Reasoning in Highly Reactive Environments. As reasoning in highly reactive environments, we identify the setting in which a knowledge-based agent, with given goals, is deployed in an environment subject…

Artificial Intelligence · Computer Science 2019-09-19 Francesco Pacenza

Explanations for AI models in high-stakes domains like medicine often lack verifiability, which can hinder trust. To address this, we propose an interactive agent that produces explanations through an auditable sequence of actions. The…

Artificial Intelligence · Computer Science 2025-11-04 Yuhang Huang , Zekai Lin , Fan Zhong , Lei Liu

A key challenge for the safety of advanced AI systems is the possibility that multiple simpler agents might inadvertently form a collective agent with capabilities and goals distinct from those of any individual. More generally, determining…

Artificial Intelligence · Computer Science 2026-05-04 Frederik Hytting Jørgensen , Sebastian Weichwald , Lewis Hammond

We introduce and study the problem of detecting whether an agent is updating their prior beliefs given new evidence in an optimal way that is Bayesian, or whether they are biased towards their own prior. In our model, biased agents form…

Computer Science and Game Theory · Computer Science 2024-10-31 Yiling Chen , Tao Lin , Ariel D. Procaccia , Aaditya Ramdas , Itai Shapira

As AI systems advance beyond human capabilities, scalable oversight becomes critical: how can we supervise AI that exceeds our abilities? A key challenge is that human evaluators may form incorrect beliefs about AI behavior in complex…

Artificial Intelligence · Computer Science 2025-10-22 Leon Lang , Patrick Forré

We argue that an explainable artificial intelligence must possess a rationale for its decisions, be able to infer the purpose of observed behaviour, and be able to explain its decisions in the context of what its audience understands and…

Artificial Intelligence · Computer Science 2021-04-26 Michael Timothy Bennett , Yoshihiro Maruyama

In order to construct an ethical artificial intelligence (AI) two complex problems must be overcome. Firstly, humans do not consistently agree on what is or is not ethical. Second, contemporary AI and machine learning methods tend to be…

Artificial Intelligence · Computer Science 2024-04-30 Michael Timothy Bennett , Yoshihiro Maruyama

Explainable models in Artificial Intelligence are often employed to ensure transparency and accountability of AI systems. The fidelity of the explanations are dependent upon the algorithms used as well as on the fidelity of the data. Many…

Machine Learning · Computer Science 2019-07-31 Muhammad Aurangzeb Ahmad , Carly Eckert , Ankur Teredesai

Handling trust is one of the core requirements for facilitating effective interaction between the human and the AI agent. Thus, any decision-making framework designed to work with humans must possess the ability to estimate and leverage…

Artificial Intelligence · Computer Science 2023-01-31 Zahra Zahedi , Sarath Sreedharan , Subbarao Kambhampati

In order to have effective human-AI collaboration, it is necessary to address how the AI agent's behavior is being perceived by the humans-in-the-loop. When the agent's task plans are generated without such considerations, they may often…

Artificial Intelligence · Computer Science 2019-03-15 Anagha Kulkarni , Yantian Zha , Tathagata Chakraborti , Satya Gautam Vadlamudi , Yu Zhang , Subbarao Kambhampati

Deployed, autonomous AI systems must often evaluate multiple plausible courses of action (extended sequences of behavior) in novel or under-specified contexts. Despite extensive training, these systems will inevitably encounter scenarios…

Artificial Intelligence · Computer Science 2025-11-19 Steven J. Jones , Robert E. Wray , John E. Laird

For artificial intelligence to be beneficial to humans the behaviour of AI agents needs to be aligned with what humans want. In this paper we discuss some behavioural issues for language agents, arising from accidental misspecification by…

Artificial Intelligence · Computer Science 2021-03-30 Zachary Kenton , Tom Everitt , Laura Weidinger , Iason Gabriel , Vladimir Mikulik , Geoffrey Irving

AI agents -- systems that combine foundation models with reasoning, planning, memory, and tool use -- are rapidly becoming a practical interface between natural-language intent and real-world computation. This survey synthesizes the…

Artificial Intelligence · Computer Science 2026-01-06 Bin Xu

How to attribute responsibility for autonomous artificial intelligence (AI) systems' actions has been widely debated across the humanities and social science disciplines. This work presents two experiments ($N$=200 each) that measure…

Computers and Society · Computer Science 2021-02-02 Gabriel Lima , Nina Grgić-Hlača , Meeyoung Cha

Agents of general intelligence deployed in real-world scenarios must adapt to ever-changing environmental conditions. While such adaptive agents may leverage engineered knowledge, they will require the capacity to construct and evaluate…

Artificial Intelligence · Computer Science 2016-06-20 Craig Sherstan , Adam White , Marlos C. Machado , Patrick M. Pilarski