中文
相关论文

相关论文: Where Paths Split: Localized, Calibrated Control o…

200 篇论文

Large language models (LLMs) increasingly participate in morally sensitive decision-making, yet how they organize ethical frameworks across reasoning steps remains underexplored. We introduce \textit{moral reasoning trajectories}, sequences…

计算与语言 · 计算机科学 2026-03-18 Fan Huang , Haewoon Kwak , Jisun An

As large language models (LLMs) increasingly participate in tasks with ethical and societal stakes, a critical question arises: do they exhibit an emergent "moral mind" - a consistent structure of moral preferences guiding their decisions -…

计算机与社会 · 计算机科学 2025-04-28 Avner Seror

Large language models (LLMs) are increasingly deployed in domains requiring moral understanding, yet their reasoning often remains shallow, and misaligned with human reasoning. Unlike humans, whose moral reasoning integrates contextual…

人机交互 · 计算机科学 2025-06-19 Mohna Chakraborty , Lu Wang , David Jurgens

As large language models (LLMs) are increasingly deployed in consequential decision-making contexts, systematically assessing their ethical reasoning capabilities becomes a critical imperative. This paper introduces the Priorities in…

人工智能 · 计算机科学 2025-04-29 Chad Coleman , W. Russell Neuman , Ali Dasdan , Safinah Ali , Manan Shah

As large language models (LLMs) become more deeply integrated into various sectors, understanding how they make moral judgments has become crucial, particularly in the realm of autonomous driving. This study utilized the Moral Machine…

计算与语言 · 计算机科学 2024-02-08 Kazuhiro Takemoto

Large Language Models (LLMs) are increasingly deployed in multilingual and multicultural environments where moral reasoning is essential for generating ethically appropriate responses. Yet, the dominant pretraining of LLMs on…

计算与语言 · 计算机科学 2025-09-29 Sualeha Farid , Jayden Lin , Zean Chen , Shivani Kumar , David Jurgens

Large Language Models (LLMs) have become central to advancing automation and decision-making across various sectors, raising significant ethical questions. This study proposes a comprehensive comparative analysis of the most advanced LLMs…

人工智能 · 计算机科学 2024-06-07 Alejandro Tlaie

Large Language Models (LLMs) have shown impressive moral reasoning abilities. Yet they often diverge when confronted with complex, multi-factor moral dilemmas. To address these discrepancies, we propose a framework that synthesizes multiple…

计算与语言 · 计算机科学 2026-02-09 Chenchen Yuan , Zheyu Zhang , Shuo Yang , Bardh Prenkaj , Gjergji Kasneci

As AI systems increasingly navigate applications in healthcare, law, and governance, understanding how they handle ethically complex scenarios becomes critical. Previous work has mainly examined the moral judgments in large language models…

In this paper we propose a framework for ethical decision making in the context of planning, with intended application to robotics. We put forward a compact but highly expressive language for ethical planning that combines linear temporal…

人工智能 · 计算机科学 2022-06-03 Umberto Grandi , Emiliano Lorini , Timothy Parker , Rachid Alami

As large language models (LLMs) increasingly mediate ethically sensitive decisions, understanding their moral reasoning processes becomes imperative. This study presents a comprehensive empirical evaluation of 14 leading LLMs, both…

计算与语言 · 计算机科学 2025-08-12 Junchen Ding , Penghao Jiang , Zihao Xu , Ziqi Ding , Yichen Zhu , Jiaojiao Jiang , Yuekang Li

We localize the policy routing mechanism in alignment-trained language models. An intermediate-layer attention gate reads detected content and triggers deeper amplifier heads that boost the signal toward refusal. In smaller models the gate…

计算与语言 · 计算机科学 2026-05-04 Gregory N. Frank

Existing behavioral alignment techniques for Large Language Models (LLMs) often neglect the discrepancy between surface compliance and internal unaligned representations, leaving LLMs vulnerable to long-tail risks. More crucially, we posit…

计算与语言 · 计算机科学 2026-03-17 Lingyu Li , Yan Teng , Yingchun Wang

We present an ethical decision-making framework that refines a pre-trained reinforcement learning (RL) model using a task-agnostic ethical layer. Following initial training, the RL model undergoes ethical fine-tuning, where human feedback…

计算机与社会 · 计算机科学 2026-05-05 Rohit K. Dubey , Damian Dailisan , Sachit Mahajan

Large language models are increasingly influencing human moral decisions, yet current approaches focus primarily on evaluating rather than actively steering their moral decisions. We formulate this as an out-of-distribution moral alignment…

人工智能 · 计算机科学 2025-11-18 Zhiyu An , Wan Du

Language models often misinterpret human intentions due to their handling of ambiguity, a limitation well-recognized in NLP research. While morally clear scenarios are more discernible to LLMs, greater difficulty is encountered in morally…

计算与语言 · 计算机科学 2024-10-11 Pranav Senthilkumar , Visshwa Balasubramanian , Prisha Jain , Aneesa Maity , Jonathan Lu , Kevin Zhu

We build a custom transformer model to study how neural networks make moral decisions on trolley-style dilemmas. The model processes structured scenarios using embeddings that encode who is affected, how many people, and which outcome they…

人工智能 · 计算机科学 2026-02-05 Mayank Goel , Aritra Das , Paras Chopra

Moral foundation detection is crucial for analyzing social discourse and developing ethically-aligned AI systems. While large language models excel across diverse tasks, their performance on specialized moral reasoning remains unclear. This…

计算与语言 · 计算机科学 2025-07-25 Maciej Skorski , Alina Landowska

This study examines the ethical reasoning of six prominent generative large language models: OpenAI GPT-4o, Meta LLaMA 3.1, Perplexity, Anthropic Claude 3.5 Sonnet, Google Gemini, and Mistral 7B. The research explores how these models…

人工智能 · 计算机科学 2025-01-16 W. Russell Neuman , Chad Coleman , Manan Shah

Artificial writing is permeating our lives due to recent advances in large-scale, transformer-based language models (LMs) such as BERT, its variants, GPT-2/3, and others. Using them as pre-trained models and fine-tuning them for specific…

计算与语言 · 计算机科学 2022-02-15 Patrick Schramowski , Cigdem Turan , Nico Andersen , Constantin A. Rothkopf , Kristian Kersting
‹ 上一页 1 2 3 10 下一页 ›