English
Related papers

Related papers: Where Paths Split: Localized, Calibrated Control o…

200 papers

Large language models are increasingly being used in critical domains of politics, business, and education, but the nature of their normative ethical judgment remains opaque. Alignment research has, to date, not sufficiently utilized…

Computers and Society · Computer Science 2025-11-18 Peter Kirgis

Enhancing the moral alignment of Large Language Models (LLMs) is a critical challenge in AI safety. Current alignment techniques often act as superficial guardrails, leaving the intrinsic moral representations of LLMs largely untouched. In…

Computation and Language · Computer Science 2026-01-16 Luoming Hu , Jingjie Zeng , Liang Yang , Hongfei Lin

Ensuring that Large Language Models (LLMs) return just responses which adhere to societal values is crucial for their broader application. Prior research has shown that LLMs often fail to perform satisfactorily on tasks requiring moral…

Computation and Language · Computer Science 2025-10-09 Guangliang Liu , Zimo Qi , Xitong Zhang , Lei Jiang , Kristen Marie Johnson

Despite significant progress in alignment, large language models (LLMs) remain vulnerable to adversarial attacks that elicit harmful behaviors. Activation steering techniques offer a promising inference-time intervention approach, but…

Machine Learning · Computer Science 2026-01-28 Quy-Anh Dang , Chris Ngo

Human moral judgment is context-dependent and modulated by interpersonal relationships. As large language models (LLMs) increasingly function as decision-support systems, determining whether they encode these social nuances is critical. We…

Computation and Language · Computer Science 2026-04-24 Jiseon Kim , Jea Kwon , Luiz Felipe Vecchietti , Wenchao Dong , Jaehong Kim , Meeyoung Cha

Large Language Models (LLMs) are known to acquire reasoning capabilities through shared inference patterns in pre-training data, which are further elicited via Chain-of-Thought (CoT) practices. However, whether fundamental reasoning…

Computation and Language · Computer Science 2026-05-28 Xingwei Tan , Marco Valentino , Mahmud Elahi Akhter , Yuxiang Zhou , Maria Liakata , Nikolaos Aletras

Recent advances in Large Language Models (LLMs) highlight the need to align their behaviors with human values. A critical, yet understudied, issue is the potential divergence between an LLM's stated preferences (its reported alignment with…

Artificial Intelligence · Computer Science 2025-06-03 Zhuojun Gu , Quan Wang , Shuchu Han

With the rise and widespread use of Large Language Models (LLMs), ensuring their safety is crucial to prevent harm to humans and promote ethical behaviors. However, directly assessing value valence (i.e., support or oppose) by leveraging…

Computers and Society · Computer Science 2025-04-10 Yuxi Sun , Wei Gao , Jing Ma , Hongzhan Lin , Ziyang Luo , Wenxuan Zhang

Inference-time computation has greatly enhanced the performance of large language models (LLMs) on challenging reasoning tasks, but this strategy can incur high inference costs. One solution is to route intermediate chain-of-thought (CoT)…

Artificial Intelligence · Computer Science 2026-05-08 Wenwen Si , Insup Lee , Osbert Bastani

Pluralism alignment with AI has the sophisticated and necessary goal of creating AI that can coexist with and serve morally multifaceted humanity. Research towards pluralism alignment has many efforts in enhancing the learning of large…

Computation and Language · Computer Science 2026-04-21 Ying Su , Mingen Zheng , Weili Diao , Haoran Li

Ethical reasoning is a crucial skill for Large Language Models (LLMs). However, moral values are not universal, but rather influenced by language and culture. This paper explores how three prominent LLMs -- GPT-4, ChatGPT, and…

Computation and Language · Computer Science 2024-04-30 Utkarsh Agarwal , Kumar Tanmay , Aditi Khandelwal , Monojit Choudhury

LLM routing aims to select the most appropriate model for each query, balancing competing performance metrics such as accuracy and cost across a pool of language models. Prior approaches typically adopt a decoupled strategy, where the…

Artificial Intelligence · Computer Science 2026-01-05 Asterios Tsiourvas , Wei Sun , Georgia Perakis

Moral norms vary across cultures. A recent line of work suggests that English large language models contain human-like moral biases, but these studies typically do not examine moral variation in a diverse cultural setting. We investigate…

Computation and Language · Computer Science 2023-06-06 Aida Ramezani , Yang Xu

Recent advances in Large Language Models (LLMs) have enabled human-like responses across various tasks, raising questions about their ethical decision-making capabilities and potential biases. This study systematically evaluates how nine…

Computers and Society · Computer Science 2025-11-03 Wentao Xu , Yile Yan , Yuqi Zhu

The question of how to make decisions that maximise the well-being of all persons is very relevant to design language models that are beneficial to humanity and free from harm. We introduce the Greatest Good Benchmark to evaluate the moral…

Computation and Language · Computer Science 2025-03-26 Giovanni Franco Gabriel Marraffini , Andrés Cotton , Noe Fabian Hsueh , Axel Fridman , Juan Wisznia , Luciano Del Corro

Large Language Models (LLMs) are increasingly positioned as decision engines for hiring, healthcare, and economic judgment, yet real-world human judgment reflects a balance between rational deliberation and emotion-driven bias. If LLMs are…

The conceptual framework proposed in this paper centers on the development of a deliberative moral reasoning system - one designed to process complex moral situations by generating, filtering, and weighing normative arguments drawn from…

Computers and Society · Computer Science 2025-08-13 David-Doron Yaacov

As large language models (LLMs) increasingly act as autonomous agents in markets and organizations, their behavior in strategic environments becomes economically consequential. We document that off-the-shelf LLM agents exhibit systematic…

General Economics · Economics 2026-03-16 Wei Lu , Amit Dhanda , Daniel L. Chen , Christian B. Hansen

Recent advancements in large language models (LLMs) have established them as powerful tools across numerous domains. However, persistent concerns about embedded biases, such as gender, racial, and cultural biases arising from their training…

Computation and Language · Computer Science 2025-07-30 Hadi Mohammadi , Yasmeen F. S. S. Meijer , Efthymia Papadopoulou , Ayoub Bagheri

People increasingly rely on Large Language Models (LLMs) for moral advice, which may influence humans' decisions. Yet, little is known about how closely LLMs align with human moral judgments. To address this, we introduce the Moral Dilemma…

Computation and Language · Computer Science 2025-07-24 Giuseppe Russo , Debora Nozza , Paul Röttger , Dirk Hovy