中文
相关论文

相关论文: Does Moral Code Have a Moral Code? Probing Delphi'…

200 篇论文

Allowing machines to choose whether to kill humans would be devastating for world peace and security. But how do we equip machines with the ability to learn ethical or even moral choices? Jentzsch et al.(2019) showed that applying machine…

计算与语言 · 计算机科学 2019-12-12 Patrick Schramowski , Cigdem Turan , Sophie Jentzsch , Constantin Rothkopf , Kristian Kersting

This paper presents a theoretical framework for the AI ethical resonance hypothesis, which proposes that advanced AI systems with purposefully designed cognitive structures ("ethical resonators") may emerge with the ability to identify…

计算机与社会 · 计算机科学 2025-07-21 Tomasz Zgliczyński-Cuber

As AI systems become more advanced and widely deployed, there will likely be increasing debate over whether AI systems could have conscious experiences, desires, or other states of potential moral significance. It is important to inform…

机器学习 · 计算机科学 2023-11-16 Ethan Perez , Robert Long

Practical uses of Artificial Intelligence (AI) in the real world have demonstrated the importance of embedding moral choices into intelligent agents. They have also highlighted that defining top-down ethical constraints on AI according to…

多智能体系统 · 计算机科学 2023-08-31 Elizaveta Tennant , Stephen Hailes , Mirco Musolesi

One of the most remarkable things about the human moral mind is its flexibility. We can make moral judgments about cases we have never seen before. We can decide that pre-established rules should be broken. We can invent novel rules on the…

The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making? Much alignment research assumes that the appropriate benchmark is how humans themselves would act…

计算机与社会 · 计算机科学 2026-05-13 Benjamin Minhao Chen , Xinyu Xie

How can we build AI systems that can learn any set of individual human values both quickly and safely, avoiding causing harm or violating societal standards for acceptable behavior during the learning process? We explore the effects of…

人工智能 · 计算机科学 2024-11-11 Andrea Wynn , Ilia Sucholutsky , Thomas L. Griffiths

As generative language models are deployed in ever-wider contexts, concerns about their political values have come to the forefront with critique from all parts of the political spectrum that the models are biased and lack neutrality.…

计算机与社会 · 计算机科学 2023-06-23 Jonathan H. Rystrøm

In this work, we study the alignment (BrainScore) of large language models (LLMs) fine-tuned for moral reasoning on behavioral data and/or brain data of humans performing the same task. We also explore if fine-tuning several LLMs on the…

人工智能 · 计算机科学 2024-11-26 Artem Karpov , Seong Hah Cho , Austin Meek , Raymond Koopmanschap , Lucy Farnik , Bogdan-Ionut Cirstea

Moral foundation detection is crucial for analyzing social discourse and developing ethically-aligned AI systems. While large language models excel across diverse tasks, their performance on specialized moral reasoning remains unclear. This…

计算与语言 · 计算机科学 2025-07-25 Maciej Skorski , Alina Landowska

Large Language Models (LLMs) are increasingly deployed in multilingual and multicultural environments where moral reasoning is essential for generating ethically appropriate responses. Yet, the dominant pretraining of LLMs on…

计算与语言 · 计算机科学 2025-09-29 Sualeha Farid , Jayden Lin , Zean Chen , Shivani Kumar , David Jurgens

As artificial intelligence systems become increasingly agentic, capable of general reasoning, planning, and value prioritization, current safety practices that treat obedience as a proxy for ethical behavior are becoming inadequate. This…

人工智能 · 计算机科学 2025-07-04 Joseph Boland

Deploying large language models (LLMs) with agency in real-world applications raises critical questions about how these models will behave. In particular, how will their decisions align with humans when faced with moral dilemmas? This study…

计算机与社会 · 计算机科学 2025-04-16 Jiseon Kim , Jea Kwon , Luiz Felipe Vecchietti , Alice Oh , Meeyoung Cha

We introduce a new computational model of moral decision making, drawing on a recent theory of commonsense moral learning via social dynamics. Our model describes moral dilemmas as a utility function that computes trade-offs in values over…

人工智能 · 计算机科学 2018-01-16 Richard Kim , Max Kleiman-Weiner , Andres Abeliuk , Edmond Awad , Sohan Dsouza , Josh Tenenbaum , Iyad Rahwan

This paper adds to the efforts of evolutionary ethics to naturalize morality by providing specific insights derived from a computational ethics view. We propose a stylized model of human decision-making, which is based on Reinforcement…

计算机与社会 · 计算机科学 2023-07-24 Eduardo C. Garrido-Merchán , Sara Lumbreras-Sancho

Large language models (LLMs) exhibit expert-level performance in tasks across a wide range of different domains. Ethical issues raised by LLMs and the need to align future versions makes it important to know how state of the art models…

Explainable AI (XAI) aims to bridge the gap between complex algorithmic systems and human stakeholders. Current discourse often examines XAI in isolation as either a technological tool, user interface, or policy mechanism. This paper…

计算机与社会 · 计算机科学 2023-11-28 Joshua L. M. Brand , Luca Nannini

A growing body of work in Ethical AI attempts to capture human moral judgments through simple computational models. The key question we address in this work is whether such simple AI models capture {the critical} nuances of moral…

Large language models (LLMs) have become integral tools in diverse domains, yet their moral reasoning capabilities across cultural and linguistic contexts remain underexplored. This study investigates whether multilingual LLMs, such as…

计算与语言 · 计算机科学 2024-12-30 Meltem Aksoy

In the rapidly evolving field of artificial intelligence, large language models (LLMs) have emerged as powerful tools for a myriad of applications, from natural language processing to decision-making support systems. However, as these…

计算与语言 · 计算机科学 2025-07-08 Jianchao Ji , Yutong Chen , Mingyu Jin , Wujiang Xu , Wenyue Hua , Yongfeng Zhang