中文
相关论文

相关论文: Value Engineering for Autonomous Agents

200 篇论文

Much of the existing research on the social and ethical impact of Artificial Intelligence has been focused on defining ethical principles and guidelines surrounding Machine Learning (ML) and other Artificial Intelligence (AI) algorithms…

计算机与社会 · 计算机科学 2019-12-30 Alexandra Luccioni , Yoshua Bengio

An important step in the development of value alignment (VA) systems in AI is understanding how VA can reflect valid ethical principles. We propose that designers of VA systems incorporate ethics by utilizing a hybrid approach in which both…

人工智能 · 计算机科学 2020-12-23 Tae Wan Kim , John Hooker , Thomas Donaldson

The rise of artificial intelligence (AI) as super-capable assistants has transformed productivity and decision-making across domains. Yet, this integration raises critical concerns about value alignment - ensuring AI behaviors remain…

人工智能 · 计算机科学 2025-10-07 Santhosh Kumar Ravindran

Contemporary debates in AI ethics increasingly foreground the prospective moral status of artificial intelligence and the possibility of extending moral or legal rights to artificial agents. While such discussions raise substantive…

计算机与社会 · 计算机科学 2026-04-07 Rahulrajan Karthikeyan , Moses Boudourides

Agreement Technologies refer to open computer systems in which autonomous software agents interact with one another, typically on behalf of humans, in order to come to mutually acceptable agreements. With the advance of AI systems in recent…

计算机与社会 · 计算机科学 2026-02-05 Andrés Holgado-Sánchez , Holger Billhardt , Alberto Fernández , Sascha Ossowski

Aligning AI agents with human values is challenging due to diverse and subjective notions of values. Standard alignment methods often aggregate crowd feedback, which can result in the suppression of unique or minority preferences. We…

人工智能 · 计算机科学 2024-10-30 Carter Blair , Kate Larson , Edith Law

AI Alignment research seeks to align human and AI goals to ensure independent actions by a machine are always ethical. This paper argues empathy is necessary for this task, despite being often neglected in favor of more deductive…

神经与进化计算 · 计算机科学 2023-12-14 Devin Gonier , Adrian Adduci , Cassidy LoCascio

As AI agents become increasingly autonomous, widely deployed in consequential contexts, and efficacious in bringing about real-world impacts, ensuring that their decisions are not only instrumentally effective but also normatively aligned…

人工智能 · 计算机科学 2026-02-03 Felix Jahn , Yannic Muskalla , Lisa Dargasz , Patrick Schramowski , Kevin Baum

When we design and deploy an Reinforcement Learning (RL) agent, reward functions motivates agents to achieve an objective. An incorrect or incomplete specification of the objective can result in behavior that does not align with human…

人工智能 · 计算机科学 2024-06-03 Zhaoyue Wang

The rapid rise of AI-based autonomous agents is transforming human society and economic systems, as these entities increasingly exhibit human-like or superhuman intelligence. From excelling at complex games like Go to tackling diverse…

人工智能 · 计算机科学 2025-05-27 Ke Yang , ChengXiang Zhai

A broad current application of algorithms is in formal and quantitative measures of murky concepts -- like merit -- to make decisions. When people strategically respond to these sorts of evaluations in order to gain favorable decision…

计算机与社会 · 计算机科学 2023-10-06 Benjamin Laufer , Jon Kleinberg , Karen Levy , Helen Nissenbaum

principles that should govern autonomous AI systems. It essentially states that a system's goals and behaviour should be aligned with human values. But how to ensure value alignment? In this paper we first provide a formal model to…

人工智能 · 计算机科学 2024-02-08 Carles Sierra , Nardine Osman , Pablo Noriega , Jordi Sabater-Mir , Antoni Perelló

The value-alignment problem for artificial intelligence (AI) asks how we can ensure that the 'values' (i.e., objective functions) of artificial systems are aligned with the values of humanity. In this paper, I argue that linguistic…

人工智能 · 计算机科学 2022-07-05 Travis LaCroix

Experts in Artificial Intelligence (AI) development predict that advances in the development of intelligent systems and agents will reshape vital areas in our society. Nevertheless, if such an advance isn't done with prudence, it can result…

人工智能 · 计算机科学 2021-08-25 Nythamar de Oliveira , Nicholas Kluge Corrêa

Road vehicle travel at a reasonable speed involves some risk, even when using computer-controlled driving with failure-free hardware and perfect sensing. A fully-automated vehicle must continuously decide how to allocate this risk without a…

计算机与社会 · 计算机科学 2020-10-30 Noah J. Goodall

The AI landscape demands a broad set of legal, ethical, and societal considerations to be accounted for in order to develop ethical AI (eAI) solutions which sustain human values and rights. Currently, a variety of guidelines and a handful…

计算机与社会 · 计算机科学 2021-12-03 Anna Felländer , Jonathan Rebane , Stefan Larsson , Mattias Wiggberg , Fredrik Heintz

Recent advances in large language models (LLMs) have enabled the development of AI agents that exhibit increasingly human-like behaviors, including planning, adaptation, and social dynamics across diverse, interactive, and open-ended…

Growing concerns about safety and alignment of AI systems highlight the importance of embedding moral capabilities in artificial agents: a promising solution is the use of learning from experience, i.e., Reinforcement Learning. In…

多智能体系统 · 计算机科学 2026-02-11 Elizaveta Tennant , Stephen Hailes , Mirco Musolesi

The concepts of blameworthiness and wrongness are of fundamental importance in human moral life. But to what extent are humans disposed to blame artificially intelligent agents, and to what extent will they judge their actions to be morally…

计算机与社会 · 计算机科学 2021-02-09 Michael T. Stuart , Markus Kneer

Recent advances in large language models (LLMs) have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical alignment a key AI safety concern. While prior work has examined both…