中文
相关论文

相关论文: Where Paths Split: Localized, Calibrated Control o…

200 篇论文

In this position paper, we argue that instead of morally aligning LLMs to specific set of ethical principles, we should infuse generic ethical reasoning capabilities into them so that they can handle value pluralism at a global scale. When…

计算与语言 · 计算机科学 2023-10-12 Abhinav Rao , Aditi Khandelwal , Kumar Tanmay , Utkarsh Agarwal , Monojit Choudhury

Autonomous systems increasingly require moral judgment capabilities, yet whether these capabilities scale predictably with model size remains unexplored. We systematically evaluate 75 large language model configurations (0.27B--1000B…

计算机与社会 · 计算机科学 2026-05-04 Kazuhiro Takemoto

As large language models (LLMs) increasingly participate in high-stakes decision-making, a central societal debate has revolved around which moral frameworks-deontological or utilitarian-should guide machine behavior. However, a largely…

计算机与社会 · 计算机科学 2026-04-14 Pengzhao Lyu , Yeun Joon Kim , Yingyue Luna Luan , Jungmin Choi

This survey paper outlines the key developments in the field of Large Language Models (LLMs), including enhancements to their reasoning skills, adaptability to various tasks, increased computational efficiency, and the ability to make…

When large language models make ethical judgments, do their internal representations distinguish between normative frameworks, or collapse ethics into a single acceptability dimension? We probe hidden representations across five ethical…

计算与语言 · 计算机科学 2026-03-26 Weilun Xu , Alexander Rusnak , Frederic Kaplan

Current alignment evaluation mostly measures whether models encode dangerous concepts and whether they refuse harmful requests. Both miss the layer where alignment often operates: routing from concept detection to behavioral policy. We…

机器学习 · 计算机科学 2026-05-04 Gregory N. Frank

Deploying large language models (LLMs) with agency in real-world applications raises critical questions about how these models will behave. In particular, how will their decisions align with humans when faced with moral dilemmas? This study…

计算机与社会 · 计算机科学 2025-04-16 Jiseon Kim , Jea Kwon , Luiz Felipe Vecchietti , Alice Oh , Meeyoung Cha

Large Language Models (LLMs) are important tools for reasoning and problem-solving, while they often operate passively, answering questions without actively discovering new ones. This limitation reduces their ability to simulate human-like…

计算工程、金融与科学 · 计算机科学 2025-09-26 Hong Su

We explore how large language models (LLMs) can be influenced by prompting them to alter their initial decisions and align them with established ethical frameworks. Our study is based on two experiments designed to assess the susceptibility…

计算与语言 · 计算机科学 2024-11-19 Allison Huang , Yulu Niki Pi , Carlos Mougan

Large language models (LLMs) exhibit expert-level performance in tasks across a wide range of different domains. Ethical issues raised by LLMs and the need to align future versions makes it important to know how state of the art models…

Large language models (LLMs) have become increasingly pivotal in various domains due the recent advancements in their performance capabilities. However, concerns persist regarding biases in LLMs, including gender, racial, and cultural…

人工智能 · 计算机科学 2024-12-03 Mijntje Meijer , Hadi Mohammadi , Ayoub Bagheri

Moral reasoning is a complex cognitive process shaped by individual experiences and cultural contexts and presents unique challenges for computational analysis. While natural language processing (NLP) offers promising tools for studying…

计算与语言 · 计算机科学 2025-02-21 Shivani Kumar , David Jurgens

In this work, we study the alignment (BrainScore) of large language models (LLMs) fine-tuned for moral reasoning on behavioral data and/or brain data of humans performing the same task. We also explore if fine-tuning several LLMs on the…

人工智能 · 计算机科学 2024-11-26 Artem Karpov , Seong Hah Cho , Austin Meek , Raymond Koopmanschap , Lucy Farnik , Bogdan-Ionut Cirstea

Large Language Models (LLMs) achieve strong performance through extended inference-time deliberation, yet how their reasoning failures arise remains poorly understood. By analyzing model-generated reasoning trajectories, we find that errors…

人工智能 · 计算机科学 2026-04-17 Wei Zhu , Jian Zhang , Lixing Yu , Kun Yue , Zhiwen Tang

When LLMs judge moral dilemmas, do they reach different conclusions in different languages, and if so, why? Two factors could drive such differences: the language of the dilemma itself, or the language in which the model reasons. Standard…

计算与语言 · 计算机科学 2026-01-16 Nan Li , Bo Kang , Tijl De Bie

Ensuring that Large Language Models (LLMs) align with the diverse and evolving human values across different regions and cultures remains a critical challenge in AI ethics. Current alignment approaches often yield superficial conformity…

人工智能 · 计算机科学 2025-11-04 Jiahao Wang , Songkai Xue , Jinghui Li , Xiaozhen Wang

Large Language Models (LLMs) have shown strong performance across many tasks, but their ability to capture culturally diverse moral values remains unclear. In this paper, we examine whether LLMs mirror variations in moral attitudes reported…

计算与语言 · 计算机科学 2026-03-31 Hadi Mohammadi , Ayoub Bagheri

Big models have greatly advanced AI's ability to understand, generate, and manipulate information and content, enabling numerous applications. However, as these models become increasingly integrated into everyday life, their inherent…

计算机与社会 · 计算机科学 2023-10-27 Xiaoyuan Yi , Jing Yao , Xiting Wang , Xing Xie

As large language models (LLMs) become increasingly integrated into society, their alignment with human morals is crucial. To better understand this alignment, we created a large corpus of human- and LLM-generated responses to various moral…

人机交互 · 计算机科学 2024-10-11 Basile Garcia , Crystal Qian , Stefano Palminteri

Do large language models reason morally, or do they merely sound like they do? We investigate whether LLM responses to moral dilemmas exhibit genuine developmental progression through Kohlberg's stages of moral development, or whether…

人工智能 · 计算机科学 2026-03-24 Aryan Kasat , Smriti Singh , Aman Chadha , Vinija Jain