中文
相关论文

相关论文: Observation Interference in Partially Observable A…

200 篇论文

This study investigates malicious AI Assistants' manipulative traits and whether the behaviours of malicious AI Assistants can be detected when interacting with human-like simulated users in various decision-making contexts. We also examine…

密码学与安全 · 计算机科学 2025-04-08 Yulu Pi , Ella Bettison , Anna Becker

In many real world contexts, successful human-AI collaboration requires humans to productively integrate complementary sources of information into AI-informed decisions. However, in practice human decision-makers often lack understanding of…

人机交互 · 计算机科学 2023-01-30 Kenneth Holstein , Maria De-Arteaga , Lakshmi Tumati , Yanghuidi Cheng

Hadfield-Menell et al. (2017) propose the Off-Switch Game, a model of Human-AI cooperation in which AI agents always defer to humans because they are uncertain about our preferences. I explain two reasons why AI agents might not defer.…

人工智能 · 计算机科学 2025-02-14 Sven Neth

As AI systems grow more capable and autonomous, ensuring their safety and reliability requires not only model-level alignment but also strategic oversight of the humans and institutions involved in their development and deployment. Existing…

人工智能 · 计算机科学 2026-02-10 Cheol Woo Kim , Davin Choo , Tzeh Yuan Neoh , Milind Tambe

Stochastic games are a well established model for multi-agent sequential decision making under uncertainty. In practical applications, though, agents often have only partial observability of their environment. Furthermore, agents…

计算机科学与博弈论 · 计算机科学 2024-07-02 Rui Yan , Gabriel Santos , Gethin Norman , David Parker , Marta Kwiatkowska

As AI usage becomes more prevalent in social contexts, understanding agent-user interaction is critical to designing systems that improve both individual and group outcomes. We present an online behavioral experiment (N = 243) in which…

计算机科学与博弈论 · 计算机科学 2026-02-16 Kehang Zhu , Nithum Thain , Vivian Tsai , James Wexler , Crystal Qian

AI design characteristics and human personality traits each impact the quality and outcomes of human-AI interactions. However, their relative and joint impacts are underexplored in imperfectly cooperative scenarios, where people and AI only…

Modern AI assistants are trained to follow instructions, implicitly assuming that users can clearly articulate their goals and the kind of assistance they need. Decades of behavioral research, however, show that people often engage with AI…

人工智能 · 计算机科学 2026-04-24 Nathanael Jo , Zoe De Simone , Mitchell Gordon , Ashia Wilson

Online information ecosystems are now central to our everyday social interactions. Of the many opportunities and challenges this presents, the capacity for artificial agents to shape individual and collective human decision-making in such…

种群与进化 · 定量生物学 2023-12-07 Theodor Cimpeanu , Alexander J. Stewart

Multiagent decision-making in partially observable environments is usually modelled as either an extensive-form game (EFG) in game theory or a partially observable stochastic game (POSG) in multiagent reinforcement learning (MARL). One…

人工智能 · 计算机科学 2021-09-29 Vojtěch Kovařík , Martin Schmid , Neil Burch , Michael Bowling , Viliam Lisý

Empirical human-AI alignment aims to make AI systems act in line with observed human behavior. While noble in its goals, we argue that empirical alignment can inadvertently introduce statistical biases that warrant caution. This position…

人工智能 · 计算机科学 2025-05-13 Julian Rodemann , Esteban Garces Arias , Christoph Luther , Christoph Jansen , Thomas Augustin

The AI alignment problem, which focusses on ensuring that artificial intelligence (AI), including AGI and ASI, systems act according to human values, presents profound challenges. With the progression from narrow AI to Artificial General…

人工智能 · 计算机科学 2025-07-25 Alberto Hernández-Espinosa , Felipe S. Abrahão , Olaf Witkowski , Hector Zenil

Recent work in the behavioural sciences has begun to overturn the long-held belief that human decision making is irrational, suboptimal and subject to biases. This turn to the rational suggests that human decision making may be a better…

机器学习 · 计算机科学 2020-06-09 Haiyang Chen , Hyung Jin Chang , Andrew Howes

Corrigibility of autonomous agents is an under explored part of system design, with previous work focusing on single agent systems. It has been suggested that uncertainty over the human preferences acts to keep the agents corrigible, even…

计算机科学与博弈论 · 计算机科学 2025-01-10 Edmund Dable-Heath , Boyko Vodenicharski , James Bishop

Value alignment problems arise in scenarios where the specified objectives of an AI agent don't match the true underlying objective of its users. The problem has been widely argued to be one of the central safety problems in AI.…

人工智能 · 计算机科学 2023-02-10 Malek Mechergui , Sarath Sreedharan

Effective collaboration between humans and AIs hinges on transparent communication and alignment of mental models. However, explicit, verbal communication is not always feasible. Under such circumstances, human-human teams often depend on…

We aim to help users estimate the state of the world in tasks like robotic teleoperation and navigation with visual impairments, where users may have systematic biases that lead to suboptimal behavior: they might struggle to process…

机器学习 · 计算机科学 2020-08-10 Siddharth Reddy , Sergey Levine , Anca D. Dragan

Adaptive machines have the potential to assist or interfere with human behavior in a range of contexts, from cognitive decision-making to physical device assistance. Therefore it is critical to understand how machine learning algorithms can…

人工智能 · 计算机科学 2023-05-03 Benjamin J. Chasnov , Lillian J. Ratliff , Samuel A. Burden

Interaction and cooperation with humans are overarching aspirations of artificial intelligence (AI) research. Recent studies demonstrate that AI agents trained with deep reinforcement learning are capable of collaborating with humans. These…

人机交互 · 计算机科学 2024-05-10 Kevin R. McKee , Xuechunzi Bai , Susan T. Fiske

We study decision-making with rational inattention in settings where agents have perception constraints. In such settings, inaccurate prior beliefs or models of others may lead to inattention blindness, where an agent is unaware of its…

计算机科学与博弈论 · 计算机科学 2025-10-06 Mustafa O. Karabag , Jesse Milzman , Ufuk Topcu