中文
相关论文

相关论文: MoralityGym: A Benchmark for Evaluating Hierarchic…

200 篇论文

The question of how to make decisions that maximise the well-being of all persons is very relevant to design language models that are beneficial to humanity and free from harm. We introduce the Greatest Good Benchmark to evaluate the moral…

As large language models (LLMs) increasingly mediate ethically sensitive decisions, understanding their moral reasoning processes becomes imperative. This study presents a comprehensive empirical evaluation of 14 leading LLMs, both…

计算与语言 · 计算机科学 2025-08-12 Junchen Ding , Penghao Jiang , Zihao Xu , Ziqi Ding , Yichen Zhu , Jiaojiao Jiang , Yuekang Li

Pluralism alignment with AI has the sophisticated and necessary goal of creating AI that can coexist with and serve morally multifaceted humanity. Research towards pluralism alignment has many efforts in enhancing the learning of large…

计算与语言 · 计算机科学 2026-04-21 Ying Su , Mingen Zheng , Weili Diao , Haoran Li

Recent advances in AI research make it increasingly plausible that artificial agents with consequential real-world impact will soon operate beyond tightly controlled environments. Ensuring that these agents are not only safe but that they…

计算机与社会 · 计算机科学 2025-06-10 Kevin Baum

With the development of Large Language Models (LLMs) in consulting, their role in moral decision-making has become prominent. However, existing research predominantly consider AI as an independent "moral agent" adhering to the "Human-AI…

社会与信息网络 · 计算机科学 2026-03-24 Yangyi Wu , Tianqi Wang , Xilin Liu

While the operationalisation of high-level AI ethics principles into practical AI/ML systems has made progress, there is still a theory-practice gap in managing tensions between the underlying AI ethics aspects. We cover five approaches for…

计算机与社会 · 计算机科学 2024-12-25 Conrad Sanderson , Emma Schleiger , David Douglas , Petra Kuhnert , Qinghua Lu

The computer security research community regularly tackles ethical questions. The field of ethics / moral philosophy has for centuries considered what it means to be "morally good" or at least "morally allowed / acceptable". Among…

密码学与安全 · 计算机科学 2023-08-08 Tadayoshi Kohno , Yasemin Acar , Wulf Loh

We review practical challenges in building and deploying ethical AI at the scale of contemporary industrial and societal uses. Apart from the purely technical concerns that are the usual focus of academic research, the operational…

机器学习 · 计算机科学 2021-08-16 Jiahao Chen , Victor Storchan , Eren Kurshan

Large Language Models (LLMs) are increasingly tasked with creative generation, including the simulation of fictional characters. However, their ability to portray non-prosocial, antagonistic personas remains largely unexamined. We…

计算与语言 · 计算机科学 2025-11-13 Zihao Yi , Qingxuan Jiang , Ruotian Ma , Xingyu Chen , Qu Yang , Mengru Wang , Fanghua Ye , Ying Shen , Zhaopeng Tu , Xiaolong Li , Linus

Ethics in AI becomes a global topic of interest for both policymakers and academic researchers. In the last few years, various research organizations, lawyers, think tankers and regulatory bodies get involved in developing AI ethics…

计算机与社会 · 计算机科学 2021-09-17 Arif Ali Khan , Sher Badshah , Peng Liang , Bilal Khan , Muhammad Waseem , Mahmood Niazi , Muhammad Azeem Akbar

Work on morality in large language models (LLMs) has progressed via constitutional AI, reinforcement learning from human feedback (RLHF) and systematic benchmarking, yet it still lacks tools to connect internal moral representations to…

人机交互 · 计算机科学 2026-03-25 Gunter Bombaerts

This work contributes to the field of Machine Ethics (ME) benchmarking, which develops tests to assess whether intelligent systems accurately represent human values and act accordingly. We identify three major issues with current ME…

计算机与社会 · 计算机科学 2024-11-11 Kira Sam , Raja Vavekanand

Algorithm fairness in the application of artificial intelligence (AI) is essential for a better society. As the foundational axiom of social mechanisms, fairness consists of multiple facets. Although the machine learning (ML) community has…

机器学习 · 计算机科学 2023-01-19 Dangxing Chen , Luyao Zhang

The rise of artificial intelligence (AI) as super-capable assistants has transformed productivity and decision-making across domains. Yet, this integration raises critical concerns about value alignment - ensuring AI behaviors remain…

人工智能 · 计算机科学 2025-10-07 Santhosh Kumar Ravindran

As Artificial Intelligence (AI) becomes pervasive in most fields, from healthcare to autonomous driving, it is essential that we find successful ways of building morality into our machines, especially for decision-making. However, the…

人工智能 · 计算机科学 2023-10-13 Reneira Seeamber , Cosmin Badea

Artificial Moral Agents (AMA's) is a field in computer science with the purpose of creating autonomous machines that can make moral decisions akin to how humans do. Researchers have proposed theoretical means of creating such machines,…

计算机与社会 · 计算机科学 2020-03-03 George Rautenbach , C. Maria Keet

Organisations that design and deploy artificial intelligence (AI) systems increasingly commit themselves to high-level, ethical principles. However, there still exists a gap between principles and practices in AI ethics. One major obstacle…

计算机与社会 · 计算机科学 2024-07-09 Jakob Mokander , Margi Sheth , David Watson , Luciano Floridi

Background: What counts as violence is neither self-evident nor universally agreed upon. While physical aggression is prototypical, contemporary societies increasingly debate whether exclusion, humiliation, online harassment or symbolic…

物理与社会 · 物理学 2026-02-20 Mariachiara Stellato , Francesco Lancia , Chiara Galeazzi , Nico Curti

Recent improvements in large language model (LLM) performance on academic benchmarks, such as MATH and GSM8K, have emboldened their use as standalone tutors and as simulations of human learning. However, these new applications require more…

人工智能 · 计算机科学 2025-05-06 Daniel Weitekamp , Momin N. Siddiqui , Christopher J. MacLellan

The ethics of Machine Learning has become an unavoidable topic in the AI Community. The deployment of machine learning systems in multiple social contexts has resulted in a closer ethical scrutiny of the design, development, and application…

计算机与社会 · 计算机科学 2022-01-19 Miguel Sicart , Irina Shklovski , Mirabelle Jones