中文
相关论文

相关论文: Probing Ethical Framework Representations in Large…

200 篇论文

We demonstrate how easy it is for modern machine-learned systems to violate common deontological ethical principles and social norms such as "favor the less fortunate," and "do not penalize good attributes." We propose that in some cases…

机器学习 · 计算机科学 2020-03-16 Serena Wang , Maya Gupta

Recently, Large Language Models (LLMs) have demonstrated a superior ability to serve as ranking models. However, concerns have arisen as LLMs will exhibit discriminatory ranking behaviors based on users' sensitive attributes (\eg gender).…

信息检索 · 计算机科学 2024-09-26 Chen Xu , Wenjie Wang , Yuxin Li , Liang Pang , Jun Xu , Tat-Seng Chua

Large language models (LLMs) have exploded in popularity in the past few years and have achieved undeniably impressive results on benchmarks as varied as question answering and text summarization. We provide a simple new prompting strategy…

计算与语言 · 计算机科学 2022-12-14 Joshua Albrecht , Ellie Kitanidis , Abraham J. Fetterman

As Large Language Models increasingly mediate human communication and decision-making, understanding their value expression becomes critical for research across disciplines. This work presents the Ethics Engine, a modular Python pipeline…

计算机与社会 · 计算机科学 2025-10-15 Jake Van Clief , Constantine Kyritsopoulos

This report summarizes a short study of the performance of GPT-4 on the ETHICS dataset. The ETHICS dataset consists of five sub-datasets covering different fields of ethics: Justice, Deontology, Virtue Ethics, Utilitarianism, and…

计算与语言 · 计算机科学 2023-09-20 Sergey Rodionov , Zarathustra Amadeus Goertzel , Ben Goertzel

Understanding the latent space geometry of large language models (LLMs) is key to interpreting their behavior and improving alignment. Yet it remains unclear to what extent LLMs linearly organize representations related to semantic…

计算与语言 · 计算机科学 2026-01-22 Baturay Saglam , Paul Kassianik , Blaine Nelson , Sajana Weerawardhena , Yaron Singer , Amin Karbasi

As Large Language Models (LLMs) become more prevalent, concerns about their safety, ethics, and potential biases have risen. Systematically evaluating LLMs' risk decision-making tendencies and attitudes, particularly in the ethical domain,…

计算机与社会 · 计算机科学 2025-05-09 Yifan Zeng , Liang Kairong , Fangzhou Dong , Peijia Zheng

The growing interest in employing large language models (LLMs) for decision-making in social and economic contexts has raised questions about their potential to function as agents in these domains. A significant number of societal problems…

计算机科学与博弈论 · 计算机科学 2025-11-25 Hadi Hosseini , Samarth Khanna

When survival instincts conflict with human welfare, how do Large Language Models (LLMs) make ethical choices? This fundamental tension becomes critical as LLMs integrate into autonomous systems with real-world consequences. We introduce…

计算机与社会 · 计算机科学 2025-09-16 Alireza Mohamadi , Ali Yavari

Large language models (LLMs) have demonstrated substantial commonsense understanding through numerous benchmark evaluations. However, their understanding of cultural commonsense remains largely unexamined. In this paper, we conduct a…

计算与语言 · 计算机科学 2024-05-09 Siqi Shen , Lajanugen Logeswaran , Moontae Lee , Honglak Lee , Soujanya Poria , Rada Mihalcea

Recently, computer scientists have developed large language models (LLMs) by training prediction models with large-scale language corpora and human reinforcements. The LLMs have become one promising way to implement artificial intelligence…

计算机与社会 · 计算机科学 2023-08-22 Hyemin Han

This paper critiques common patterns in machine ethics for Reinforcement Learning (RL) and argues for a virtue focused alternative. We highlight two recurring limitations in much of the current literature: (i) rule based (deontological)…

人工智能 · 计算机科学 2026-04-08 Majid Ghasemi , Mark Crowley

Massively multilingual sentence representations are trained on large corpora of uncurated data, with a very imbalanced proportion of languages included in the training. This may cause the models to grasp cultural values including moral…

Recent advances in NLP show that language models retain a discernible level of knowledge in deontological ethics and moral norms. However, existing works often treat morality as binary, ranging from right to wrong. This simplistic view does…

计算与语言 · 计算机科学 2024-05-28 Jeongwoo Park , Enrico Liscio , Pradeep K. Murukannaiah

The rapid advancement and adaptability of Large Language Models (LLMs) highlight the need for moral consistency, the capacity to maintain ethically coherent reasoning across varied contexts. Existing alignment frameworks, structured…

计算与语言 · 计算机科学 2025-12-03 Saeid Jamshidi , Kawser Wazed Nafi , Arghavan Moradi Dakhel , Negar Shahabi , Foutse Khomh

Natural Language Processing prides itself to be an empirically-minded, if not outright empiricist field, and yet lately it seems to get itself into essentialist debates on issues of meaning and measurement ("Do Large Language Models…

计算与语言 · 计算机科学 2023-10-30 David Schlangen

As AI systems become an increasing part of people's everyday lives, it becomes ever more important that they understand people's ethical norms. Motivated by descriptive ethics, a field of study that focuses on people's descriptive judgments…

计算与语言 · 计算机科学 2021-03-25 Nicholas Lourie , Ronan Le Bras , Yejin Choi

Aligning large language models (LLMs) with a human reasoning approach ensures that LLMs produce morally correct and human-like decisions. Ethical concerns are raised because current models are prone to generating false positives and…

Understanding what defines a good representation in large language models (LLMs) is fundamental to both theoretical understanding and practical applications. In this paper, we investigate the quality of intermediate representations in…

机器学习 · 计算机科学 2024-12-13 Oscar Skean , Md Rifat Arefin , Yann LeCun , Ravid Shwartz-Ziv

Despite recent advancements showcasing the impressive capabilities of Large Language Models (LLMs) in conversational systems, we show that even state-of-the-art LLMs are morally inconsistent in their generations, questioning their…

计算与语言 · 计算机科学 2024-03-11 Vamshi Krishna Bonagiri , Sreeram Vennam , Priyanshul Govil , Ponnurangam Kumaraguru , Manas Gaur