中文
相关论文

相关论文: Learning the Value Systems of Societies from Prefe…

200 篇论文

There has been considerable interest in predicting human emotions and traits using facial images and videos. Lately, such work has come under criticism for poor labeling practices, inconclusive prediction results and fairness…

计算机视觉与模式识别 · 计算机科学 2020-07-13 Abhishek Singhania , Abhishek Unnam , Varun Aggarwal

Humans strive to design safe AI systems that align with our goals and remain under our control. However, as AI capabilities advance, we face a new challenge: the emergence of deeper, more persistent relationships between humans and AI…

人机交互 · 计算机科学 2025-02-05 Hannah Rose Kirk , Iason Gabriel , Chris Summerfield , Bertie Vidgen , Scott A. Hale

Many modern machine learning approaches require vast amounts of training data to learn new concepts; conversely, human learning often requires few examples--sometimes only one--from which the learner can abstract structural concepts. We…

人工智能 · 计算机科学 2018-11-28 Nikhil Krishnaswamy , Scott Friedman , James Pustejovsky

This paper examines the challenge of embedding public values into national artificial intelligence (AI) governance frameworks, a task complicated by the sociotechnical nature of contemporary systems. As AI permeates domains such as…

计算机与社会 · 计算机科学 2026-02-19 Mike Wa Nkongolo

AI Alignment research seeks to align human and AI goals to ensure independent actions by a machine are always ethical. This paper argues empathy is necessary for this task, despite being often neglected in favor of more deductive…

神经与进化计算 · 计算机科学 2023-12-14 Devin Gonier , Adrian Adduci , Cassidy LoCascio

We present an overview of the literature on trust in AI and AI trustworthiness and argue for the need to distinguish these concepts more clearly and to gather more empirically evidence on what contributes to people s trusting behaviours. We…

人工智能 · 计算机科学 2023-09-20 Andreas Duenser , David M. Douglas

As artificial intelligence (AI) continues to advance--particularly in generative models--an open question is whether these systems can replicate foundational models of human social perception. A well-established framework in social…

计算与语言 · 计算机科学 2025-03-10 Necdet Gurkan , Kimathi Njoki , Jordan W. Suchow

Rational agents are usually built to maximize rewards. However, AGI agents can find undesirable ways of maximizing any prior reward function. Therefore value learning is crucial for safe AGI. We assume that generalized states of the world…

人工智能 · 计算机科学 2013-08-06 Alexey Potapov , Sergey Rodionov

Effective coordination and cooperation among agents are crucial for accomplishing individual or shared objectives in multi-agent systems. In many real-world multi-agent systems, agents possess varying abilities and constraints, making it…

多智能体系统 · 计算机科学 2023-10-20 Yasin Findik , Paul Robinette , Kshitij Jerath , S. Reza Ahmadzadeh

The rise of Artificial Intelligence (AI) will bring with it an ever-increasing willingness to cede decision-making to machines. But rather than just giving machines the power to make decisions that affect us, we need ways to work…

计算机与社会 · 计算机科学 2020-12-14 Elisa Bertino , Finale Doshi-Velez , Maria Gini , Daniel Lopresti , David Parkes

Algorithms learned from data are increasingly used for deciding many aspects in our life: from movies we see, to prices we pay, or medicine we get. Yet there is growing evidence that decision making by inappropriately trained algorithms may…

人工智能 · 计算机科学 2017-08-03 Indre Zliobaite

"Human-aware" has become a popular keyword used to describe a particular class of AI systems that are designed to work and interact with humans. While there exists a surprising level of consistency among the works that use the label…

机器人学 · 计算机科学 2024-07-16 Silvia Tulli , Stylianos Loukas Vasileiou , Sarath Sreedharan

AI is being increasingly used to aid response efforts to humanitarian emergencies at multiple levels of decision-making. Such AI systems are generally understood to be stand-alone tools for decision support, with ethical assessments,…

计算机与社会 · 计算机科学 2022-09-23 Joseph Aylett-Bullock , Miguel Luengo-Oroz

Over the next few years, society as a whole will need to address what core values it wishes to protect when dealing with technology. Anthropology, a field dedicated to the very notion of what it means to be human, can provide some…

计算机与社会 · 计算机科学 2020-10-08 Alexandrine Royer

This paper looks at philosophical questions that arise in the context of AI alignment. It defends three propositions. First, normative and technical aspects of the AI alignment problem are interrelated, creating space for productive…

计算机与社会 · 计算机科学 2020-10-07 Iason Gabriel

AI has revolutionised decision-making across various fields. Yet human judgement remains paramount for high-stakes decision-making. This has fueled explorations of collaborative decision-making between humans and AI systems, aiming to…

人机交互 · 计算机科学 2026-01-22 Simran Kaur , Sara Salimzadeh , Ujwal Gadiraju

As generative AI models become increasingly integrated into high-stakes domains, the need for robust methods to evaluate their ethical reasoning becomes increasingly important. This paper introduces a five-dimensional audit model --…

人工智能 · 计算机科学 2025-04-25 W. Russell Neuman , Chad Coleman , Ali Dasdan , Safinah Ali , Manan Shah

Reward modelling from preference data is a crucial step in aligning large language models (LLMs) with human values, requiring robust generalisation to novel prompt-response pairs. In this work, we propose to frame this problem in a causal…

人工智能 · 计算机科学 2026-05-12 Katarzyna Kobalczyk , Mihaela van der Schaar

Decision making in crucial applications such as lending, hiring, and college admissions has witnessed increasing use of algorithmic models and techniques as a result of a confluence of factors such as ubiquitous connectivity, ability to…

人工智能 · 计算机科学 2020-09-08 G Roshan Lal , Sahin Cem Geyik , Krishnaram Kenthapadi

Modelling human variation in rating tasks is crucial for personalization, pluralistic model alignment, and computational social science. We propose representing individuals using natural language value profiles -- descriptions of underlying…