中文
相关论文

相关论文: Social Choice Should Guide AI Alignment in Dealing…

200 篇论文

Human feedback can prevent overtly harmful utterances in conversational models, but may not automatically mitigate subtle problematic behaviors such as a stated desire for self-preservation or power. Constitutional AI offers an alternative,…

With the growing adoption of AI systems, reasoning about how society can exert control over AI becomes an increasingly urgent problem. Existing work on democratic control largely focuses on macro-level governance. In contrast, we propose a…

计算机与社会 · 计算机科学 2026-05-19 Paul Anton Bachmann , Niclas Boehmer , Lukas Daniel Klausner , Martin Lackner

Social choice theory is the study of preference aggregation across a population, used both in mechanism design for human agents and in the democratic alignment of language models. In this study, we propose the representative social choice…

机器学习 · 计算机科学 2025-11-03 Tianyi Qiu

Human feedback is commonly utilized to finetune AI assistants. But human feedback may also encourage model responses that match user beliefs over truthful ones, a behaviour known as sycophancy. We investigate the prevalence of sycophancy in…

Given that Artificial Intelligence (AI) increasingly permeates our lives, it is critical that we systematically align AI objectives with the goals and values of humans. The human-AI alignment problem stems from the impracticality of…

计算机与社会 · 计算机科学 2022-07-05 John Nay , James Daily

The mathematical study of voting, social choice theory, has traditionally only been applicable to choices among a few predetermined alternatives, but not to open-ended decisions such as collectively selecting a textual statement. We…

计算机科学与博弈论 · 计算机科学 2025-03-07 Sara Fish , Paul Gölz , David C. Parkes , Ariel D. Procaccia , Gili Rusak , Itai Shapira , Manuel Wüthrich

This paper investigates the voting behaviors of Large Language Models (LLMs), specifically GPT-4 and LLaMA-2, their biases, and how they align with human voting patterns. Our methodology involved using a dataset from a human voting…

计算与语言 · 计算机科学 2024-12-20 Joshua C. Yang , Damian Dailisan , Marcin Korecki , Carina I. Hausladen , Dirk Helbing

Artificial Intelligence has the potential to exacerbate societal bias and set back decades of advances in equal rights and civil liberty. Data used to train machine learning algorithms may capture social injustices, inequality or…

计算机与社会 · 计算机科学 2020-08-18 Susan Leavy , Barry O'Sullivan , Eugenia Siapera

Recent work on the limitations of using reinforcement learning from human feedback (RLHF) to incorporate human preferences into model behavior often raises social choice theory as a reference point. Social choice theory's analysis of…

人工智能 · 计算机科学 2024-04-22 Jessica Dai , Eve Fleisig

Innovations in AI have focused primarily on the questions of "what" and "how"-algorithms for finding patterns in web searches, for instance-without adequate attention to the possible harms (such as privacy, bias, or manipulation) and…

计算机与社会 · 计算机科学 2020-12-14 Suresh Venkatasubramanian , Nadya Bliss , Helen Nissenbaum , Melanie Moses

A crucial consideration when developing and deploying Large Language Models (LLMs) is the human values to which these models are aligned. In the constitutional framework of alignment models are aligned to a set of principles (the…

机器学习 · 计算机科学 2026-01-27 Henry Bell , Lara Neubauer da Costa Schertel , Bochu Ding , Brandon Fain

AI systems are often used to make or contribute to important decisions in a growing range of applications, including criminal justice, hiring, and medicine. Since these decisions impact human lives, it is important that the AI systems act…

Current AI systems minimize risk by enforcing ideological neutrality, yet this may introduce automation bias by suppressing cognitive engagement in human decision-making. We conducted randomized trials with 2,500 participants to test…

人机交互 · 计算机科学 2025-08-21 Shiyang Lai , Junsol Kim , Nadav Kunievsky , Yujin Potter , James Evans

Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models. Pairwise text preferences, where human or AI annotators select the "better" of two options, are particularly common. Such preferences are used to train…

计算与语言 · 计算机科学 2025-04-22 Arduin Findeis , Timo Kaufmann , Eyke Hüllermeier , Samuel Albanie , Robert Mullins

Culture fundamentally shapes people's reasoning, behavior, and communication. As people increasingly use generative artificial intelligence (AI) to expedite and automate personal and professional tasks, cultural values embedded in AI models…

计算与语言 · 计算机科学 2024-09-20 Yan Tao , Olga Viberg , Ryan S. Baker , Rene F. Kizilcec

Preference tuning is a crucial process for aligning deep generative models with human preferences. This survey offers a thorough overview of recent advancements in preference tuning and the integration of human feedback. The paper is…

计算与语言 · 计算机科学 2024-11-05 Genta Indra Winata , Hanyang Zhao , Anirban Das , Wenpin Tang , David D. Yao , Shi-Xiong Zhang , Sambit Sahu

Autoregressive language models, which use deep learning to produce human-like texts, have become increasingly widespread. Such models are powering popular virtual assistants in areas like smart health, finance, and autonomous driving. While…

人工智能 · 计算机科学 2023-03-16 Kaiping Chen , Anqi Shao , Jirayu Burapacheep , Yixuan Li

Customising AI technologies to each user's preferences is fundamental to them functioning well. Unfortunately, current methods require too much user involvement and fail to capture their true preferences. In fact, to avoid the nuisance of…

信息检索 · 计算机科学 2023-08-14 Marc Serramia , Natalia Criado , Michael Luck

There is growing consensus that language model (LM) developers should not be the sole deciders of LM behavior, creating a need for methods that enable the broader public to collectively shape the behavior of LM systems that affect them. To…

人工智能 · 计算机科学 2024-06-13 Saffron Huang , Divya Siddarth , Liane Lovitt , Thomas I. Liao , Esin Durmus , Alex Tamkin , Deep Ganguli

Generative artificial intelligence (AI) holds enormous potential to revolutionize decision-making processes, from everyday to high-stake scenarios. By leveraging generative AI, humans can benefit from data-driven insights and predictions,…

综合经济学 · 经济学 2024-02-19 Valerio Capraro , Roberto Di Paolo , Veronica Pizziol
‹ 上一页 1 2 3 10 下一页 ›