English
Related papers

Related papers: Probing Ethical Framework Representations in Large…

200 papers

Moral judgment is integral to large language models' (LLMs) social reasoning. As multi-agent systems gain prominence, it becomes crucial to understand how LLMs function when collaborating compared to operating as individual agents. In human…

Computation and Language · Computer Science 2025-10-30 Anita Keshmirian , Razan Baltaji , Babak Hemmatian , Hadi Asghari , Lav R. Varshney

Although behavioral studies have documented numerical reasoning errors in large language models (LLMs), the underlying representational mechanisms remain unclear. We hypothesize that numerical attributes occupy shared latent subspaces and…

Artificial Intelligence · Computer Science 2025-11-11 Hirohane Takagi , Gouki Minegishi , Shota Kizawa , Issey Sukeda , Hitomi Yanaka

Developing AI systems capable of nuanced ethical reasoning is critical as they increasingly influence human decisions, yet existing models often rely on superficial correlations rather than principled moral understanding. This paper…

Computers and Society · Computer Science 2025-10-16 Mahamodul Hasan Mahadi , Md. Nasif Safwan , Souhardo Rahman , Shahnaj Parvin , Aminun Nahar , Kamruddin Nur

Large Language Model safety alignment predominantly operates on a binary assumption that requests are either safe or unsafe. This classification proves insufficient when models encounter ethical dilemmas, where the capacity to reason…

Cryptography and Security · Computer Science 2026-04-16 Shei Pern Chua , Zhen Leng Thai , Kai Jun Teh , Xiao Li , Qibing Ren , Xiaolin Hu

The rapid proliferation of Large Language Models (LLMs) has raised significant trustworthiness and ethical concerns. Despite the widespread adoption of LLMs across domains, there is still no clear consensus on how to define and…

Computation and Language · Computer Science 2025-05-06 José Siqueira de Cerqueira , Kai-Kristian Kemell , Rebekah Rousi , Nannan Xi , Juho Hamari , Pekka Abrahamsson

Critical decision-making in socially consequential spaces is increasingly involving AI systems at varying capacities. Yet, despite the ubiquity of autonomous systems, most approaches to handling autonomous moral decision-making resort to…

Artificial Intelligence · Computer Science 2026-05-28 Aisha Aijaz , Rahul Goel , Arnav Batra , Raghava Mutharaju

Reliable simulation of human behavior is essential for explaining, predicting, and intervening in our society. Recent advances in large language models (LLMs) have shown promise in emulating human behaviors, interactions, and…

Computation and Language · Computer Science 2025-10-27 Ning Bian , Xianpei Han , Hongyu Lin , Baolei Wu , Jun Wang

We present a study on how and where personas -- defined by distinct sets of human characteristics, values, and beliefs -- are encoded in the representation space of large language models (LLMs). Using a range of dimension reduction and…

Computation and Language · Computer Science 2025-09-10 Celia Cintas , Miriam Rateike , Erik Miehling , Elizabeth Daly , Skyler Speakman

Large language models generate judgments that resemble those of humans. Yet the extent to which these models align with human judgments in interpreting figurative and socially grounded language remains uncertain. To investigate this, human…

Computation and Language · Computer Science 2026-01-15 Samhita Bollepally , Aurora Sloman-Moll , Takashi Yamauchi

While existing evaluations of large language models (LLMs) measure deception rates, the underlying conditions that give rise to deceptive behavior are poorly understood. We investigate this question using a novel dataset of realistic moral…

We show how to assess a language model's knowledge of basic concepts of morality. We introduce the ETHICS dataset, a new benchmark that spans concepts in justice, well-being, duties, virtues, and commonsense morality. Models predict…

Computers and Society · Computer Science 2023-02-20 Dan Hendrycks , Collin Burns , Steven Basart , Andrew Critch , Jerry Li , Dawn Song , Jacob Steinhardt

Large language models (LLMs) closely interact with humans, and thus need an intimate understanding of the cultural values of human society. In this paper, we explore how open-source LLMs make judgments on diverse categories of cultural…

Computation and Language · Computer Science 2024-12-13 Minsang Kim , Seungjun Baek

Large language models (LLMs) have been actively applied in the mental health field. Recent research shows the promise of LLMs in applying psychotherapy, especially motivational interviewing (MI). However, there is a lack of studies…

Computation and Language · Computer Science 2025-04-01 Haein Kong , Seonghyeon Moon

Large language models (LLMs) are increasingly deployed in politically sensitive settings, raising concerns about their potential to encode, amplify, or be steered toward specific ideologies. We investigate how adopting synthetic personas…

Computation and Language · Computer Science 2025-08-25 Pietro Bernardelle , Stefano Civelli , Leon Fröhling , Riccardo Lunardi , Kevin Roitero , Gianluca Demartini

Large language models (LLMs) increasingly operate in social contexts, motivating analysis of how they express and shift moral judgments. In this work, we investigate the moral response of LLMs to persona role-play, prompting a LLM to assume…

Computation and Language · Computer Science 2026-05-15 Davi Bastos Costa , Felippe Alves , Renato Vicente

Large Language Models (LLMs) have shown impressive moral reasoning abilities. Yet they often diverge when confronted with complex, multi-factor moral dilemmas. To address these discrepancies, we propose a framework that synthesizes multiple…

Computation and Language · Computer Science 2026-02-09 Chenchen Yuan , Zheyu Zhang , Shuo Yang , Bardh Prenkaj , Gjergji Kasneci

Large Language Models (LLMs) have impressive capabilities, but are prone to outputting falsehoods. Recent work has developed techniques for inferring whether a LLM is telling the truth by training probes on the LLM's internal activations.…

Artificial Intelligence · Computer Science 2024-08-20 Samuel Marks , Max Tegmark

Public leaderboards increasingly suggest that large language models (LLMs) surpass human experts on benchmarks spanning academic knowledge, law, and programming. Yet most benchmarks are fully public, their questions widely mirrored across…

Artificial Intelligence · Computer Science 2026-03-18 Eshwar Reddy M , Sourav Karmakar

This paper investigates the ethical implications of aligning Large Language Models (LLMs) with financial optimization, through the case study of GreedLlama, a model fine-tuned to prioritize economically beneficial outcomes. By comparing…

Computation and Language · Computer Science 2024-04-05 Jeffy Yu , Maximilian Huber , Kevin Tang

Large Language Models significantly influence social interactions, decision-making, and information dissemination, underscoring the need to understand the implicit socio-cognitive attitudes, referred to as "worldviews", encoded within these…

Computation and Language · Computer Science 2025-12-30 Jiatao Li , Yanheng Li , Xiaojun Wan