中文
相关论文

相关论文: The dangers in algorithms learning humans' values …

200 篇论文

Technological advances of virtually every kind pose risks to society including fairness and bias. We review a long-standing wisdom that a widespread practical deployment of any technology may produce adverse side effects misusing the…

计算机与社会 · 计算机科学 2020-10-26 Simon Kasif

Assuming humans are (approximately) rational enables robots to infer reward functions by observing human behavior. But people exhibit a wide array of irrationalities, and our goal with this work is to better understand the effect they can…

机器学习 · 计算机科学 2021-11-16 Lawrence Chan , Andrew Critch , Anca Dragan

Artificial Intelligence (AI) has made impressive progress in recent years and represents a key technology that has a crucial impact on the economy and society. However, it is clear that AI and business models based on it can only reach…

Despite rapid technological progress, effective human-machine cooperation remains a significant challenge. Humans tend to cooperate less with machines than with fellow humans, a phenomenon known as the machine penalty. Here, we show that…

人机交互 · 计算机科学 2025-05-29 Zhen Wang , Ruiqi Song , Chen Shen , Shiya Yin , Zhao Song , Balaraju Battu , Lei Shi , Danyang Jia , Talal Rahwan , Shuyue Hu

Recent work in the behavioural sciences has begun to overturn the long-held belief that human decision making is irrational, suboptimal and subject to biases. This turn to the rational suggests that human decision making may be a better…

机器学习 · 计算机科学 2020-06-09 Haiyang Chen , Hyung Jin Chang , Andrew Howes

How can we ensure that AI systems are aligned with human values and remain safe? We can study this problem through the frameworks of the AI assistance and the AI shutdown games. The AI assistance problem concerns designing an AI agent that…

人工智能 · 计算机科学 2025-12-30 Alessio Benavoli , Alessandro Facchini , Marco Zaffalon

The field of AI alignment aims to steer AI systems toward human goals, preferences, and ethical principles. Its contributions have been instrumental for improving the output quality, safety, and trustworthiness of today's AI models. This…

人工智能 · 计算机科学 2024-11-26 Robert West , Roland Aydin

Can machines think? This is a central question in artificial intelligence research. However, there is a substantial divergence of views on the answer to this question. Why do people have such significant differences of opinion, even when…

人工智能 · 计算机科学 2025-12-01 Xi Cun , Jifan Ren , Asha Huang , Siyu Li , Ruzhen Song

In many real-life settings, algorithms play the role of assistants, while humans ultimately make the final decision. Often, algorithms specifically act as curators, narrowing down a wide range of options into a smaller subset that the human…

计算机科学与博弈论 · 计算机科学 2025-11-06 Jiaxin Song , Parnian Shahkar , Kate Donahue , Bhaskar Ray Chaudhury

Artificial intelligence (AI) systems can cause harm to people. This research examines how individuals react to such harm through the lens of blame. Building upon research suggesting that people blame AI systems, we investigated how several…

计算机与社会 · 计算机科学 2023-04-06 Gabriel Lima , Nina Grgić-Hlača , Meeyoung Cha

One of the major challenges we face with ethical AI today is developing computational systems whose reasoning and behaviour are provably aligned with human values. Human values, however, are notorious for being ambiguous, contradictory and…

人工智能 · 计算机科学 2023-05-05 Nardine Osman , Mark d'Inverno

The value-alignment problem for artificial intelligence (AI) asks how we can ensure that the 'values' (i.e., objective functions) of artificial systems are aligned with the values of humanity. In this paper, I argue that linguistic…

人工智能 · 计算机科学 2022-07-05 Travis LaCroix

To interact with humans, artificial intelligence (AI) systems must understand our social world. Within this world norms play an important role in motivating and guiding agents. However, very few computational theories for learning social…

人工智能 · 计算机科学 2022-01-27 Taylor Olson , Ken Forbus

Big models have greatly advanced AI's ability to understand, generate, and manipulate information and content, enabling numerous applications. However, as these models become increasingly integrated into everyday life, their inherent…

计算机与社会 · 计算机科学 2023-10-27 Xiaoyuan Yi , Jing Yao , Xiting Wang , Xing Xie

Conversational AI is rapidly becoming a primary interface for information seeking and decision making, yet most systems still assume idealized users. In practice, human reasoning is bounded by limited attention, uneven knowledge, and…

新兴技术 · 计算机科学 2026-01-21 Jiqun Liu

Artificial intelligence (AI) was initially developed as an implicit moral agent to solve simple and clearly defined tasks where all options are predictable. However, it is now part of our daily life powering cell phones, cameras, watches,…

计算机与社会 · 计算机科学 2020-02-11 Mohamed Akrout , Robert Steinbauer

AI systems are often used to make or contribute to important decisions in a growing range of applications, including criminal justice, hiring, and medicine. Since these decisions impact human lives, it is important that the AI systems act…

. It is typically assumed that for the successful use of machine learning algorithms, these algorithms should have a higher accuracy than a human expert. Moreover, if the average accuracy of ML algorithms is lower than that of a human…

人机交互 · 计算机科学 2024-11-19 Saveli Goldberg , Lev Salnikov , Noor Kaiser , Tushar Srivastava , Eugene Pinsky

How to attribute responsibility for autonomous artificial intelligence (AI) systems' actions has been widely debated across the humanities and social science disciplines. This work presents two experiments ($N$=200 each) that measure…

计算机与社会 · 计算机科学 2021-02-02 Gabriel Lima , Nina Grgić-Hlača , Meeyoung Cha

When working with generative artificial intelligence (AI), users may see productivity gains, but the AI-generated content may not match their preferences exactly. To study this effect, we introduce a Bayesian framework in which…

人工智能 · 计算机科学 2025-07-08 Francisco Castro , Jian Gao , Sébastien Martin