English
Related papers

Related papers: What are human values, and how do we align AI to t…

200 papers

Aligning AI agents to human intentions and values is a key bottleneck in building safe and deployable AI applications. But whose values should AI agents be aligned with? Reinforcement learning with human feedback (RLHF) has emerged as the…

Artificial Intelligence · Computer Science 2023-10-25 Abhilash Mishra

Benchmarks are seen as the cornerstone for measuring technical progress in Artificial Intelligence (AI) research and have been developed for a variety of tasks ranging from question answering to facial recognition. An increasingly prominent…

Computers and Society · Computer Science 2022-04-12 Travis LaCroix , Alexandra Sasha Luccioni

Aligning AI systems with human values and the value-based preferences of various stakeholders (their value systems) is key in ethical AI. In value-aware AI systems, decision-making draws upon explicit computational representations of…

Artificial Intelligence · Computer Science 2025-07-29 Andrés Holgado-Sánchez , Holger Billhardt , Sascha Ossowski , Sara Degli-Esposti

Operationalizing human values alongside functional and adaptation requirements remains challenging due to their ambiguous, pluralistic, and context-dependent nature. Explicit representations are needed to support the elicitation, analysis,…

Software Engineering · Computer Science 2026-02-11 Everaldo Silva Júnior , Lina Marsso , Ricardo Caldas , Marsha Chechik , Genaína Nunes Rodrigues

Large language models (LLMs) have demonstrated great potential for automating the evaluation of natural language generation. Previous frameworks of LLM-as-a-judge fall short in two ways: they either use zero-shot setting without consulting…

Computation and Language · Computer Science 2025-04-11 Mingxuan Li , Hanchen Li , Chenhao Tan

An important aspect of developing LLMs that interact with humans is to align models' behavior to their users. It is possible to prompt an LLM into behaving as a certain persona, especially a user group or ideological persona the model…

Computation and Language · Computer Science 2023-05-25 EunJeong Hwang , Bodhisattwa Prasad Majumder , Niket Tandon

Current AI training methods align models with human values only after their core capabilities have been established, resulting in models that are easily misaligned and lack deep-rooted value systems. We propose a paradigm shift from "model…

Artificial Intelligence · Computer Science 2025-11-18 Roland Aydin , Christian Cyron , Steve Bachelor , Ashton Anderson , Robert West

Value alignment of Large Language Models (LLMs) requires us to empirically measure these models' actual, acquired representation of value. Among the characteristics of value representation in humans is that they distinguish among value of…

Computation and Language · Computer Science 2026-02-24 Seong Hah Cho , Junyi Li , Anna Leshinskaya

Constructing a universal moral code for artificial intelligence (AI) is difficult or even impossible, given that different human cultures have different definitions of morality and different societal norms. We therefore argue that the value…

As AI systems become pervasive, grounding their behavior in human values is critical. Prior work suggests that language models (LMs) exhibit limited inherent moral reasoning, leading to calls for explicit moral teaching. However,…

Computation and Language · Computer Science 2026-01-27 Meysam Alizadeh , Fabrizio Gilardi , Zeynab Samei

Generative AI systems (ChatGPT, DALL-E, etc) are expanding into multiple areas of our lives, from art Rombach et al. [2021] to mental health Rob Morris and Kareem Kouddous [2022]; their rapidly growing societal impact opens new…

Artificial Intelligence · Computer Science 2024-05-21 Alexis Roger , Esma Aïmeur , Irina Rish

Value alignment problems arise in scenarios where the specified objectives of an AI agent don't match the true underlying objective of its users. The problem has been widely argued to be one of the central safety problems in AI.…

Artificial Intelligence · Computer Science 2023-02-10 Malek Mechergui , Sarath Sreedharan

As large language models (LLMs) enter the mainstream, aligning them to foster constructive dialogue rather than exacerbate societal divisions is critical. Using an individualized and multicultural alignment dataset of over 7,500…

Human-Computer Interaction · Computer Science 2025-03-24 Yara Kyrychenko , Jon Roozenbeek , Brandon Davidson , Sander van der Linden , Ramit Debnath

Recent advancements in the field of natural language generation have facilitated the use of large language models to assess the quality of generated text. Although these models have shown promising results in tasks such as machine…

Artificial Intelligence · Computer Science 2024-01-23 Terry Yue Zhuo

While large language models (LLMs) have been used for automated grading, they have not yet achieved the same level of performance as humans, especially when it comes to grading complex questions. Existing research on this topic focuses on a…

Artificial Intelligence · Computer Science 2024-05-31 Wenjing Xie , Juxin Niu , Chun Jason Xue , Nan Guan

Deploying large language models (LLMs) with agency in real-world applications raises critical questions about how these models will behave. In particular, how will their decisions align with humans when faced with moral dilemmas? This study…

Computers and Society · Computer Science 2025-04-16 Jiseon Kim , Jea Kwon , Luiz Felipe Vecchietti , Alice Oh , Meeyoung Cha

Recent advancements in Large Language Models (LLMs) have revolutionized the AI field but also pose potential safety and ethical risks. Deciphering LLMs' embedded values becomes crucial for assessing and mitigating their risks. Despite…

Computation and Language · Computer Science 2024-05-13 Pablo Biedma , Xiaoyuan Yi , Linus Huang , Maosong Sun , Xing Xie

AI intent alignment, ensuring that AI produces outcomes as intended by users, is a critical challenge in human-AI interaction. The emergence of generative AI, including LLMs, has intensified the significance of this problem, as interactions…

Human-Computer Interaction · Computer Science 2024-06-21 Yoonsu Kim , Kihoon Son , Seoyoung Kim , Juho Kim

Given that AI systems are set to play a pivotal role in future decision-making processes, their trustworthiness and reliability are of critical concern. Due to their scale and complexity, modern AI systems resist direct interpretation, and…

Artificial Intelligence · Computer Science 2025-01-03 Binxia Xu , Antonis Bikakis , Daniel Onah , Andreas Vlachidis , Luke Dickens

Making moral judgments is an essential step toward developing ethical AI systems. Prevalent approaches are mostly implemented in a bottom-up manner, which uses a large set of annotated data to train models based on crowd-sourced opinions…

Computation and Language · Computer Science 2024-07-02 Jingyan Zhou , Minda Hu , Junan Li , Xiaoying Zhang , Xixin Wu , Irwin King , Helen Meng