中文
相关论文

相关论文: Off-Switching Not Guaranteed

200 篇论文

To reduce the danger of powerful super-intelligent AIs, we might make the first such AIs oracles that can only send and receive messages. This paper proposes a possibly practical means of using machine learning to create two classes of…

人工智能 · 计算机科学 2020-10-07 James D. Miller , Roman Yampolskiy , Olle Haggstrom , Stuart Armstrong

AI systems increasingly support human decision-making. In many cases, despite the algorithm's superior performance, the final decision remains in human hands. For example, an AI may assist doctors in determining which diagnostic tests to…

人工智能 · 计算机科学 2026-02-20 Gali Noti , Kate Donahue , Jon Kleinberg , Sigal Oren

In some agent designs like inverse reinforcement learning an agent needs to learn its own reward function. Learning the reward function and optimising for it are typically two different processes, usually performed at different stages. We…

人工智能 · 计算机科学 2020-04-29 Stuart Armstrong , Jan Leike , Laurent Orseau , Shane Legg

While AI systems have equaled or surpassed human performance in a wide variety of games such as Chess, Go, or Dota 2, describing these systems as truly "human-like" remains far-fetched. Despite their success, they fail to replicate the…

人工智能 · 计算机科学 2025-07-09 Aloïs Rautureau , Éric Piette

Coordination and cooperation between humans and autonomous agents in cooperative games raises interesting questions of human decision making and behaviour changes. Here we report our findings from a group formation game in a small-world…

物理与社会 · 物理学 2021-05-21 Tuomas Takko , Kunal Bhattacharya , Daniel Monsivais , Kimmo Kaski

With humans interacting with AI-based systems at an increasing rate, it is necessary to ensure the artificial systems are acting in a manner which reflects understanding of the human. In the case of humans and artificial AI agents operating…

人机交互 · 计算机科学 2023-02-03 Andrew Fuchs , Andrea Passarella , Marco Conti

AI assistance in decision-making has become popular, yet people's inappropriate reliance on AI often leads to unsatisfactory human-AI collaboration performance. In this paper, through three pre-registered, randomized human subject…

人机交互 · 计算机科学 2024-01-17 Zhuoran Lu , Dakuo Wang , Ming Yin

The recent adoption of machine learning as a tool in real world decision making has spurred interest in understanding how these decisions are being made. Counterfactual Explanations are a popular interpretable machine learning technique…

机器学习 · 计算机科学 2021-10-05 Andrew O'Brien , Edward Kim

Artificial intelligence (AI) tools such as large language models (LLMs) are already altering student learning. Unlike previous technologies, LLMs can independently solve problems regardless of student understanding, yet are not always…

理论经济学 · 经济学 2025-09-04 Eric Gao

This study explores the dynamics of trust in artificial intelligence (AI) agents, particularly large language models (LLMs), by introducing the concept of "deferred trust", a cognitive mechanism where distrust in human agents redirects…

人机交互 · 计算机科学 2025-11-24 Johan Sebastián Galindez-Acosta , Juan José Giraldo-Huertas

Human-AI collaboration (HAIC) in decision-making aims to create synergistic teaming between human decision-makers and AI systems. Learning to defer (L2D) has been presented as a promising framework to determine who among humans and AI…

机器学习 · 计算机科学 2022-07-14 Diogo Leitão , Pedro Saleiro , Mário A. T. Figueiredo , Pedro Bizarro

Today, AI is increasingly being used in many high-stakes decision-making applications in which fairness is an important concern. Already, there are many examples of AI being biased and making questionable and unfair decisions. The AI…

人工智能 · 计算机科学 2020-02-06 Yunfeng Zhang , Rachel K. E. Bellamy , Kush R. Varshney

Artificial intelligence (AI) is supposed to help us make better choices. Some of these choices are small, like what route to take to work, or what music to listen to. Others are big, like what treatment to administer for a disease or how…

计算机与社会 · 计算机科学 2021-05-18 Bryce Goodman

How to attribute responsibility for autonomous artificial intelligence (AI) systems' actions has been widely debated across the humanities and social science disciplines. This work presents two experiments ($N$=200 each) that measure…

计算机与社会 · 计算机科学 2021-02-02 Gabriel Lima , Nina Grgić-Hlača , Meeyoung Cha

In the last few decades, numerous experiments have shown that humans do not always behave so as to maximize their material payoff. Cooperative behavior when non-cooperation is a dominant strategy (with respect to the material payoffs) is…

计算机科学与博弈论 · 计算机科学 2016-06-27 Valerio Capraro , Joseph Y. Halpern

In the context of humans operating with artificial or autonomous agents in a hybrid team, it is essential to accurately identify when to authorize those team members to perform actions. Given past examples where humans and autonomous…

人工智能 · 计算机科学 2023-10-12 Andrew Fuchs , Andrea Passarella , Marco Conti

Reward function, as an incentive representation that recognizes humans' agency and rationalizes humans' actions, is particularly appealing for modeling human behavior in human-robot interaction. Inverse Reinforcement Learning is an…

人工智能 · 计算机科学 2021-03-09 Ran Tian , Masayoshi Tomizuka , Liting Sun

The AI-alignment problem arises when there is a discrepancy between the goals that a human designer specifies to an AI learner and a potential catastrophic outcome that does not reflect what the human designer really wants. We argue that a…

机器学习 · 计算机科学 2020-04-10 Shai Shalev-Shwartz , Shaked Shammah , Amnon Shashua

We anticipate increased instances of humans and AI systems working together in what we refer to as a hybrid team. The increase in collaboration is expected as AI systems gain proficiency and their adoption becomes more widespread. However,…

人工智能 · 计算机科学 2024-08-06 Andrew Fuchs , Andrea Passarella , Marco Conti

Despite AI's superhuman performance in a variety of domains, humans are often unwilling to adopt AI systems. The lack of interpretability inherent in many modern AI techniques is believed to be hurting their adoption, as users may not trust…

人工智能 · 计算机科学 2021-11-17 Daehwan Ahn , Abdullah Almaatouq , Monisha Gulabani , Kartik Hosanagar