中文
相关论文

相关论文: A Multi-Level Framework for the AI Alignment Probl…

200 篇论文

Benchmarks are seen as the cornerstone for measuring technical progress in Artificial Intelligence (AI) research and have been developed for a variety of tasks ranging from question answering to facial recognition. An increasingly prominent…

计算机与社会 · 计算机科学 2022-04-12 Travis LaCroix , Alexandra Sasha Luccioni

Are AI systems truly representing human values, or merely averaging across them? Our study suggests a concerning reality: Large Language Models (LLMs) fail to represent diverse cultural moral frameworks despite their linguistic…

计算与语言 · 计算机科学 2025-08-01 Simon Münker

Research on fairness, accountability, transparency and ethics of AI-based interventions in society has gained much-needed momentum in recent years. However it lacks an explicit alignment with a set of normative values and principles that…

人工智能 · 计算机科学 2022-10-07 Vinodkumar Prabhakaran , Margaret Mitchell , Timnit Gebru , Iason Gabriel

Conversational agents (CAs) based on generative artificial intelligence frequently face challenges ensuring ethical interactions that align with human values. Current value alignment efforts largely rely on top-down approaches, such as…

计算机与社会 · 计算机科学 2025-07-30 Lenart Motnikar , Katharina Baum , Alexander Kagan , Sarah Spiekermann-Hoff

AI Alignment, primarily in the form of Reinforcement Learning from Human Feedback (RLHF), has been a cornerstone of the post-training phase in developing Large Language Models (LLMs). It has also been a popular research topic across various…

计算与语言 · 计算机科学 2025-08-26 Ilias Chalkidis

AI ethics is an emerging field with multiple, competing narratives about how to best solve the problem of building human values into machines. Two major approaches are focused on bias and compliance, respectively. But neither of these ideas…

人工智能 · 计算机科学 2023-02-24 Thomas Krendl Gilbert , Megan Welle Brozek , Andrew Brozek

The development of sophisticated artificial intelligence (AI) conversational agents based on large language models raises important questions about the relationship between human norms, values, and practices and AI design and performance.…

计算机与社会 · 计算机科学 2025-05-30 Rachel Katharine Sterken , James Ravi Kirkpatrick

As AI systems become increasingly capable and influential, ensuring their alignment with human values, preferences, and goals has become a critical research focus. Current alignment methods primarily focus on designing algorithms and loss…

计算与语言 · 计算机科学 2025-05-02 Min-Hsuan Yeh , Jeffrey Wang , Xuefeng Du , Seongheon Park , Leitian Tao , Shawn Im , Yixuan Li

As the deployment of artificial intelligence (AI) is changing many fields and industries, there are concerns about AI systems making decisions and recommendations without adequately considering various ethical aspects, such as…

计算机与社会 · 计算机科学 2023-10-02 Conrad Sanderson , Qinghua Lu , David Douglas , Xiwei Xu , Liming Zhu , Jon Whittle

This paper critically evaluates the attempts to align Artificial Intelligence (AI) systems, especially Large Language Models (LLMs), with human values and intentions through Reinforcement Learning from Feedback (RLxF) methods, involving…

The popularisation of applying AI in businesses poses significant challenges relating to ethical principles, governance, and legal compliance. Although businesses have embedded AI into their day-to-day processes, they lack a unified…

人工智能 · 计算机科学 2024-12-09 Haocheng Lin

Artificial intelligence (AI) is advancing at a pace that raises urgent questions about how to align machine decision-making with human moral values. This working paper investigates how leading AI systems prioritize moral outcomes and what…

人工智能 · 计算机科学 2025-09-15 Eoin O'Doherty , Nicole Weinrauch , Andrew Talone , Uri Klempner , Xiaoyuan Yi , Xing Xie , Yi Zeng

Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psychology's focus on mental illness: necessary but incomplete.…

The AI alignment problem comprises both technical and normative dimensions. While technical solutions focus on implementing normative constraints in AI systems, the normative problem concerns determining what these constraints should be.…

计算机与社会 · 计算机科学 2025-07-29 André Steingrüber , Kevin Baum

This position paper argues that effectively "democratizing AI" requires democratic governance and alignment of AI, and that this is particularly valuable for decisions with systemic societal impacts. Initial steps -- such as Meta's…

Instruction-tuned Large Language Models (LLMs) are increasingly deployed as AI Assistants in firms for support in cognitive tasks. These AI assistants carry embedded perspectives which influence factors across the firm including…

计算机与社会 · 计算机科学 2025-05-27 Noah Broestl , Benjamin Lange , Cristina Voinea , Geoff Keeling , Rachael Lam

One of the major challenges we face with ethical AI today is developing computational systems whose reasoning and behaviour are provably aligned with human values. Human values, however, are notorious for being ambiguous, contradictory and…

人工智能 · 计算机科学 2023-05-05 Nardine Osman , Mark d'Inverno

Pluralistic alignment is concerned with ensuring that an AI system's objectives and behaviors are in harmony with the diversity of human values and perspectives. In this paper we study the notion of pluralistic alignment in the context of…

人工智能 · 计算机科学 2024-11-19 Parand A. Alamdari , Toryn Q. Klassen , Rodrigo Toro Icarte , Sheila A. McIlraith

Artificial Intelligence (AI) applications are being used to predict and assess behaviour in multiple domains, such as criminal justice and consumer finance, which directly affect human well-being. However, if AI is to improve people's…

其他计算机科学 · 计算机科学 2019-06-12 Andrea Aler Tubella , Andreas Theodorou , Virginia Dignum , Frank Dignum

With AI systems becoming more powerful and pervasive, there is increasing debate about keeping their actions aligned with the broader goals and needs of humanity. This multi-disciplinary and multi-stakeholder debate must resolve many…

人工智能 · 计算机科学 2021-12-21 Koen Holtman