English
Related papers

Related papers: Exploring and steering the moral compass of Large …

200 papers

Large language models (LLMs) have rapidly shifted from peripheral assistive tools to constant companions in everyday and even high stakes human decision making. Many users now consult these models about health, intimate relationships,…

Computers and Society · Computer Science 2026-02-26 Abas Bertina , Sara Shakeri

Large language models (LLMs) demonstrate outstanding capabilities, but challenges remain regarding their ability to solve complex reasoning tasks, as well as their transparency, robustness, truthfulness, and ethical alignment. In this…

Computers and Society · Computer Science 2023-12-19 Konstantin Hebenstreit , Robert Praas , Matthias Samwald

In this work, we study the alignment (BrainScore) of large language models (LLMs) fine-tuned for moral reasoning on behavioral data and/or brain data of humans performing the same task. We also explore if fine-tuning several LLMs on the…

Artificial Intelligence · Computer Science 2024-11-26 Artem Karpov , Seong Hah Cho , Austin Meek , Raymond Koopmanschap , Lucy Farnik , Bogdan-Ionut Cirstea

While Large Language Models (LLMs) have become ubiquitous in many fields, understanding and mitigating LLM biases is an ongoing issue. This paper provides a novel method for evaluating the demographic biases of various generative AI models.…

Computation and Language · Computer Science 2025-06-16 Jack H Fagan , Ruhaan Juyaal , Amy Yue-Ming Yu , Siya Pun

Large Language Models (LLMs) have gained significant popularity for their application in various everyday tasks such as text generation, summarization, and information retrieval. As the widespread adoption of LLMs continues to surge, it…

Computation and Language · Computer Science 2024-03-22 Pagnarasmey Pit , Xingjun Ma , Mike Conway , Qingyu Chen , James Bailey , Henry Pit , Putrasmey Keo , Watey Diep , Yu-Gang Jiang

In this study, we measure the moral reasoning ability of LLMs using the Defining Issues Test - a psychometric instrument developed for measuring the moral development stage of a person according to the Kohlberg's Cognitive Moral Development…

Computation and Language · Computer Science 2023-10-10 Kumar Tanmay , Aditi Khandelwal , Utkarsh Agarwal , Monojit Choudhury

Large Language Models (LLMs) have made unprecedented breakthroughs, yet their increasing integration into everyday life might raise societal risks due to generated unethical content. Despite extensive study on specific issues like bias, the…

Computation and Language · Computer Science 2024-03-05 Shitong Duan , Xiaoyuan Yi , Peng Zhang , Tun Lu , Xing Xie , Ning Gu

Autonomous systems increasingly require moral judgment capabilities, yet whether these capabilities scale predictably with model size remains unexplored. We systematically evaluate 75 large language model configurations (0.27B--1000B…

Computers and Society · Computer Science 2026-05-04 Kazuhiro Takemoto

Although large language models (LLMs) demonstrate impressive proficiency in various tasks, they present potential safety risks, such as `jailbreaks', where malicious inputs can coerce LLMs into generating harmful content bypassing safety…

Computation and Language · Computer Science 2025-11-26 Isack Lee , Haebin Seong

Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "unbiased." This binary framing misses the gradual,…

Machine Learning · Computer Science 2026-05-06 Yash Aggarwal , Atmika Gorti , Vinija Jain , Aman Chadha , Krishnaprasad Thirunarayan , Manas Gaur

Large language models (LLMs), despite their remarkable capabilities, are susceptible to generating biased and discriminatory responses. As LLMs increasingly influence high-stakes decision-making (e.g., hiring and healthcare), mitigating…

Computation and Language · Computer Science 2025-03-04 Jingling Li , Zeyu Tang , Xiaoyu Liu , Peter Spirtes , Kun Zhang , Liu Leqi , Yang Liu

Work on morality in large language models (LLMs) has progressed via constitutional AI, reinforcement learning from human feedback (RLHF) and systematic benchmarking, yet it still lacks tools to connect internal moral representations to…

Human-Computer Interaction · Computer Science 2026-03-25 Gunter Bombaerts

Controlling the behavior of large language models (LLMs) at inference time is essential for aligning outputs with human abilities and safety requirements. \emph{Activation steering} provides a lightweight alternative to prompt engineering…

Artificial Intelligence · Computer Science 2026-01-30 Diaoulé Diallo , Katharina Dworatzyk , Sophie Jentzsch , Peer Schütt , Sabine Theis , Tobias Hecking

The open-sourcing of large language models (LLMs) accelerates application development, innovation, and scientific progress. This includes both base models, which are pre-trained on extensive datasets without alignment, and aligned models,…

Computation and Language · Computer Science 2024-04-17 Xiao Wang , Tianze Chen , Xianjun Yang , Qi Zhang , Xun Zhao , Dahua Lin

Large Language Models (LLMs) such as ChatGPT, have gained significant attention due to their impressive natural language processing capabilities. It is crucial to prioritize human-centered principles when utilizing these models.…

Computation and Language · Computer Science 2023-06-21 Yue Huang , Qihui Zhang , Philip S. Y , Lichao Sun

Large Language Models (LLMs) are increasingly employed in software engineering tasks such as requirements elicitation, design, and evaluation, raising critical questions regarding their alignment with human judgments on responsible AI…

Software Engineering · Computer Science 2025-11-07 Asma Yamani , Malak Baslyman , Moataz Ahmed

Large Language Models (LLMs) have made substantial progress in the past several months, shattering state-of-the-art benchmarks in many domains. This paper investigates LLMs' behavior with respect to gender stereotypes, a known issue for…

Computation and Language · Computer Science 2023-08-30 Hadas Kotek , Rikker Dockum , David Q. Sun

Large Language Models (LLMs) have emerged as powerful candidates to inform clinical decision-making processes. While these models play an increasingly prominent role in shaping the digital landscape, two growing concerns emerge in…

Computation and Language · Computer Science 2024-04-24 Raphael Poulain , Hamed Fayyaz , Rahmatollah Beheshti

Large language models (LLMs) are now deployed at unprecedented scale, assisting millions of users in daily tasks. However, the risk of these models assisting unlawful activities remains underexplored. In this study, we define this high-risk…

Computers and Society · Computer Science 2025-11-27 Xing Wang , Huiyuan Xie , Yiyan Wang , Chaojun Xiao , Huimin Chen , Holli Sargeant , Felix Steffek , Jie Shao , Zhiyuan Liu , Maosong Sun

Work in AI ethics and fairness has made much progress in regulating LLMs to reflect certain values, such as fairness, truth, and diversity. However, it has taken the problem of how LLMs might 'mean' anything at all for granted. Without…

Computation and Language · Computer Science 2023-11-07 Mark Pock , Andre Ye , Jared Moore