中文
相关论文

相关论文: Human irrationality: both bad and good for reward …

200 篇论文

We often assume that robots which collaborate with humans should behave in ways that are transparent (e.g., legible, explainable). These transparent robots intentionally choose actions that convey their internal state to nearby humans: for…

机器人学 · 计算机科学 2025-05-19 Shahabedin Sagheb , Soham Gandhi , Dylan P. Losey

The black-box nature of neural models has motivated a line of research that aims to generate natural language rationales to explain why a model made certain predictions. Such rationale generation models, to date, have been trained on…

计算与语言 · 计算机科学 2020-12-16 Faeze Brahman , Vered Shwartz , Rachel Rudinger , Yejin Choi

Imitation is a key component of human social behavior, and is widely used by both children and adults as a way to navigate uncertain or unfamiliar situations. But in an environment populated by multiple heterogeneous agents pursuing…

神经元与认知 · 定量生物学 2023-05-15 Max Taylor-Davies , Stephanie Droop , Christopher G. Lucas

Reward learning algorithms utilize human feedback to infer a reward function, which is then used to train an AI system. This human feedback is often a preference comparison, in which the human teacher compares several samples of AI behavior…

机器学习 · 计算机科学 2023-03-03 Peter Barnett , Rachel Freedman , Justin Svegliato , Stuart Russell

It is known that recommendations of AI-based systems can be incorrect or unfair. Hence, it is often proposed that a human be the final decision-maker. Prior work has argued that explanations are an essential pathway to help human…

人机交互 · 计算机科学 2022-05-10 Jakob Schoeffer , Maria De-Arteaga , Niklas Kuehl

Rationality is often related to optimal decision making. Humans are known to be bounded rational agents. However, recent advances in computing, and other scientific and technical fields along with large amount of data have led to a feeling…

计算机与社会 · 计算机科学 2023-06-21 Dibakar Das

A central component of rational behavior is logical inference: the process of determining which conclusions follow from a set of premises. Psychologists have documented several ways in which humans' inferences deviate from the rules of…

计算与语言 · 计算机科学 2024-04-12 Tiwalayo Eisape , MH Tessler , Ishita Dasgupta , Fei Sha , Sjoerd van Steenkiste , Tal Linzen

The effectiveness of reinforcement learning (RL) agents in continuous control robotics tasks is mainly dependent on the design of the underlying reward function, which is highly prone to reward hacking. A misalignment between the reward…

Is it possible to evaluate the moral cognition of complex artificial agents? In this work, we take a look at one aspect of morality: `doing the right thing for the right reasons.' We propose a behavior-based analysis of artificial moral…

A recent approach based on Bayesian inverse planning for the "theory of mind" has shown good performance in modeling human cognition. However, perfect inverse planning differs from human cognition during one kind of complex tasks due to…

人工智能 · 计算机科学 2019-11-21 Ryo Nakahashi , Seiji Yamada

Intrinsic rewards were introduced to simulate how human intelligence works; they are usually evaluated by intrinsically-motivated play, i.e., playing games without extrinsic rewards but evaluated with extrinsic rewards. However, none of the…

人工智能 · 计算机科学 2019-11-28 Yuhang Song , Jianyi Wang , Thomas Lukasiewicz , Zhenghua Xu , Shangtong Zhang , Andrzej Wojcicki , Mai Xu

To facilitate effective human-robot interaction (HRI), trust-aware HRI has been proposed, wherein the robotic agent explicitly considers the human's trust during its planning and decision making. The success of trust-aware HRI depends on…

机器人学 · 计算机科学 2021-03-19 Yaohui Guo , Cong Shi , X. Jessie Yang

As robots become increasingly prevalent in human environments, there will inevitably be times when a robot needs to interrupt a human to initiate an interaction. Our work introduces the first interruptibility-aware mobile robot system, and…

机器人学 · 计算机科学 2018-04-18 Siddhartha Banerjee , Andrew Silva , Karen Feigh , Sonia Chernova

Reward models play a critical role in guiding large language models toward outputs that align with human expectations. However, an open challenge remains in effectively utilizing test-time compute to enhance reward model performance. In…

计算与语言 · 计算机科学 2025-05-21 Jiaxin Guo , Zewen Chi , Li Dong , Qingxiu Dong , Xun Wu , Shaohan Huang , Furu Wei

Training language models with rationales augmentation has been shown to be beneficial in many existing works. In this paper, we identify that such a prevailing view does not hold consistently. We conduct comprehensive investigations to…

计算与语言 · 计算机科学 2025-06-02 Chiwei Zhu , Benfeng Xu , An Yang , Junyang Lin , Quan Wang , Chang Zhou , Zhendong Mao

The rapid development of artificial intelligence and robotics has had a significant impact on our lives, with intelligent systems increasingly performing tasks traditionally performed by humans. Efficient knowledge transfer requires…

机器人学 · 计算机科学 2025-01-10 Phillip Richter , Heiko Wersing , Anna-Lisa Vollmer

Debiasing methods in NLP models traditionally focus on isolating information related to a sensitive attribute (e.g., gender or race). We instead argue that a favorable debiasing method should use sensitive information 'fairly,' with…

计算与语言 · 计算机科学 2023-10-24 Bodhisattwa Prasad Majumder , Zexue He , Julian McAuley

Inferential decision-making algorithms typically assume that an underlying probabilistic model of decision alternatives and outcomes may be learned a priori or online. Furthermore, when applied to robots in real-world settings they often…

机器人学 · 计算机科学 2023-09-15 Yucheng Chen , Pingping Zhu , Anthony Alers , Tobias Egner , Marc A. Sommer , Silvia Ferrari

Reward learning enables the application of reinforcement learning (RL) to tasks where reward is defined by human judgment, building a model of reward by asking humans questions. Most work on reward learning has used simulated environments,…

Large language models (LLMs) have the potential to aid and improve human decision-making in classification tasks, not only by providing fairly accurate predictions, but also in their ability to generate cogent narrative explanations of…

人机交互 · 计算机科学 2026-05-25 Laura R. Marusich , Mary Grace Kozuch Dhooghe , Jonathan Z. Bakdash , Murat Kantarcioglu