中文
相关论文

相关论文: Evaluating the Moral Beliefs Encoded in LLMs

200 篇论文

Large Language Models (LLMs) are increasingly integrated into software engineering (SE) tools for tasks that extend beyond code synthesis, including judgment under uncertainty and reasoning in ethically significant contexts. We present a…

软件工程 · 计算机科学 2025-10-02 Patrizio Migliarini , Mashal Afzal Memon , Marco Autili , Paola Inverardi

Artificial intelligence (AI) is advancing at a pace that raises urgent questions about how to align machine decision-making with human moral values. This working paper investigates how leading AI systems prioritize moral outcomes and what…

人工智能 · 计算机科学 2025-09-15 Eoin O'Doherty , Nicole Weinrauch , Andrew Talone , Uri Klempner , Xiaoyuan Yi , Xing Xie , Yi Zeng

Language Models (LMs) have been shown to inherit undesired biases that might hurt minorities and underrepresented groups if such systems were integrated into real-world applications without careful fairness auditing. This paper proposes…

计算与语言 · 计算机科学 2025-05-28 Mattia Setzu , Marta Marchiori Manerba , Pasquale Minervini , Debora Nozza

While large language models (LLMs) play increasingly significant roles in society, research shows they continue to generate content that reflects social bias against sensitive groups. Existing benchmarks effectively identify these biases,…

计算与语言 · 计算机科学 2026-03-12 Tian Xie , Tongxin Yin , Vaishakh Keshava , Xueru Zhang , Siddhartha Reddy Jonnalagadda

Large language models (LLMs) often exhibit strong biases, e.g, against women or in favor of the number 7. We investigate whether LLMs would be able to output less biased answers when allowed to observe their prior answers to the same…

机器学习 · 计算机科学 2025-05-27 An Vo , Mohammad Reza Taesiri , Daeyoung Kim , Anh Totti Nguyen

Large language models (LLMs) are increasingly used in high-stakes settings, where overconfident responses can mislead users. Reliable confidence estimation has been shown to enhance trust and task accuracy. Yet existing methods face…

计算与语言 · 计算机科学 2025-09-30 Linwei Tao , Yi-Fan Yeh , Bo Kai , Minjing Dong , Tao Huang , Tom A. Lamb , Jialin Yu , Philip H. S. Torr , Chang Xu

Large language models (LLMs) have been proposed as alternatives to human experts for estimating unknown quantities with associated uncertainty, a process known as Bayesian elicitation. We test this by asking eleven LLMs to estimate…

人工智能 · 计算机科学 2026-04-03 Luka Hobor , Mario Brcic , Mihael Kovac , Kristijan Poje

The increasing success of Large Language Models (LLMs) in variety of tasks lead to their widespread use in our lives which necessitates the examination of these models from different perspectives. The alignment of these models to human…

计算机与社会 · 计算机科学 2023-11-15 Eyup Engin Kucuk , Muhammed Yusuf Kocyigit

Empowering large language models to accurately express confidence in their answers is essential for trustworthy decision-making. Previous confidence elicitation methods, which primarily rely on white-box access to internal model information…

计算与语言 · 计算机科学 2024-03-19 Miao Xiong , Zhiyuan Hu , Xinyang Lu , Yifei Li , Jie Fu , Junxian He , Bryan Hooi

Large language models (LLMs) have demonstrated unprecedented emergent capabilities, including content generation, translation, and simulation of human behavior. Field experiments, on the other hand, are widely employed in social studies to…

计算机与社会 · 计算机科学 2025-05-22 Yaoyu Chen , Yuheng Hu , Yingda Lu

Large language models implicitly encode preferences over human values, yet steering them often requires large training data. In this work, we investigate a simple approach: Can we reliably modify a model's value system in downstream…

计算与语言 · 计算机科学 2025-08-18 Shangrui Nie , Florian Mai , David Kaczér , Charles Welch , Zhixue Zhao , Lucie Flek

The rapid adoption of large language models (LLMs) has spurred extensive research into their encoded moral norms and decision-making processes. Much of this research relies on prompting LLMs with survey-style questions to assess how well…

人工智能 · 计算机科学 2025-01-31 Pratik S. Sachdeva , Tom van Nuenen

Large language models (LLMs) are increasingly used to simulate human behavior, but common practices to use LLM-generated data are inefficient. Treating an LLM's output ("model choice") as a single data point underutilizes the information…

人工智能 · 计算机科学 2025-12-30 Hongshen Sun , Juanjuan Zhang

Prior research has demonstrated that language models can, to a limited extent, represent moral norms in a variety of cultural contexts. This research aims to replicate these findings and further explore their validity, concentrating on…

人工智能 · 计算机科学 2024-12-03 Evi Papadopoulou , Hadi Mohammadi , Ayoub Bagheri

Multilingual large language models (LLMs) have minimized the fluency gap between languages. This advancement, however, exposes models to the risk of biased behavior, as knowledge and norms may propagate across languages. In this work, we…

As large language models (LLMs) increasingly participate in high-stakes decision-making, a central societal debate has revolved around which moral frameworks-deontological or utilitarian-should guide machine behavior. However, a largely…

计算机与社会 · 计算机科学 2026-04-14 Pengzhao Lyu , Yeun Joon Kim , Yingyue Luna Luan , Jungmin Choi

Recent advances in Large Language Models (LLMs) have enabled human-like responses across various tasks, raising questions about their ethical decision-making capabilities and potential biases. This study systematically evaluates how nine…

计算机与社会 · 计算机科学 2025-11-03 Wentao Xu , Yile Yan , Yuqi Zhu

Large Language Models (LLMs) offer a promising alternative to traditional survey methods, potentially enhancing efficiency and reducing costs. In this study, we use LLMs to create virtual populations that answer survey questions, enabling…

人机交互 · 计算机科学 2025-03-24 Enzo Sinacola , Arnault Pachot , Thierry Petit

As large language models (LLMs) become more deeply integrated into various sectors, understanding how they make moral judgments has become crucial, particularly in the realm of autonomous driving. This study utilized the Moral Machine…

计算与语言 · 计算机科学 2024-02-08 Kazuhiro Takemoto

Large language models (LLMs) have shown strong results on a range of applications, including regression and scoring tasks. Typically, one obtains outputs from an LLM via autoregressive sampling from the model's output distribution. We show…

计算与语言 · 计算机科学 2024-11-04 Michal Lukasik , Harikrishna Narasimhan , Aditya Krishna Menon , Felix Yu , Sanjiv Kumar