English
Related papers

Related papers: Value Imprint: A Technique for Auditing the Human …

200 papers

Reinforcement learning from human feedback (RLHF) has emerged as an effective approach to aligning large language models (LLMs) to human preferences. RLHF contains three steps, i.e., human preference collecting, reward learning, and policy…

Computation and Language · Computer Science 2024-03-29 Hao Lang , Fei Huang , Yongbin Li

Large language model (LLM) agents with extended autonomy unlock new capabilities, but also introduce heightened challenges for LLM safety. In particular, an LLM agent may pursue objectives that deviate from human values and ethical norms, a…

Computation and Language · Computer Science 2026-01-27 Chen Chen , Kim Young Il , Yuan Yang , Wenhao Su , Yilin Zhang , Xueluan Gong , Qian Wang , Yongsen Zheng , Ziyao Liu , Kwok-Yan Lam

Are AI systems truly representing human values, or merely averaging across them? Our study suggests a concerning reality: Large Language Models (LLMs) fail to represent diverse cultural moral frameworks despite their linguistic…

Computation and Language · Computer Science 2025-08-01 Simon Münker

Research on fairness, accountability, transparency and ethics of AI-based interventions in society has gained much-needed momentum in recent years. However it lacks an explicit alignment with a set of normative values and principles that…

Artificial Intelligence · Computer Science 2022-10-07 Vinodkumar Prabhakaran , Margaret Mitchell , Timnit Gebru , Iason Gabriel

The proliferation of large language models (LLMs) requires robust evaluation of their alignment with local values and ethical standards, especially as existing benchmarks often reflect the cultural, legal, and ideological values of their…

Computers and Society · Computer Science 2024-08-06 Gwenyth Isobel Meadows , Nicholas Wai Long Lau , Eva Adelina Susanto , Chi Lok Yu , Aditya Paul

Human preference judgments are pivotal in guiding large language models (LLMs) to produce outputs that align with human values. Human evaluations are also used in summarization tasks to compare outputs from various systems, complementing…

Computation and Language · Computer Science 2023-10-31 Yebowen Hu , Kaiqiang Song , Sangwoo Cho , Xiaoyang Wang , Hassan Foroosh , Fei Liu

Reinforcement learning from human feedback (RLHF) is a variant of reinforcement learning (RL) that learns from human feedback instead of relying on an engineered reward function. Building on prior work on the related setting of…

Machine Learning · Computer Science 2025-12-30 Timo Kaufmann , Paul Weng , Viktor Bengs , Eyke Hüllermeier

Rapid integration of large language models (LLMs) into societal applications has intensified concerns about their alignment with universal ethical principles, as their internal value representations remain opaque despite behavioral…

Computation and Language · Computer Science 2025-05-26 Yi Su , Jiayi Zhang , Shu Yang , Xinhai Wang , Lijie Hu , Di Wang

Reinforcement learning with human feedback (RLHF) is an emerging paradigm to align models with human preferences. Typically, RLHF aggregates preferences from multiple individuals who have diverse viewpoints that may conflict with each…

Machine Learning · Computer Science 2024-03-11 Huiying Zhong , Zhun Deng , Weijie J. Su , Zhiwei Steven Wu , Linjun Zhang

Reinforcement Learning (RL) has emerged as a transformative approach for aligning and enhancing Large Language Models (LLMs), addressing critical challenges in instruction following, ethical alignment, and reasoning capabilities. This…

Artificial Intelligence · Computer Science 2025-07-08 Saksham Sahai Srivastava , Vaneet Aggarwal

ChatGLM is a free-to-use AI service powered by the ChatGLM family of large language models (LLMs). In this paper, we present the ChatGLM-RLHF pipeline -- a reinforcement learning from human feedback (RLHF) system -- designed to enhance…

Computation and Language · Computer Science 2024-04-04 Zhenyu Hou , Yilin Niu , Zhengxiao Du , Xiaohan Zhang , Xiao Liu , Aohan Zeng , Qinkai Zheng , Minlie Huang , Hongning Wang , Jie Tang , Yuxiao Dong

The reward model for Reinforcement Learning from Human Feedback (RLHF) has proven effective in fine-tuning Large Language Models (LLMs). Notably, collecting human feedback for RLHF can be resource-intensive and lead to scalability issues…

Computation and Language · Computer Science 2024-07-09 Jinghan Zhang , Xiting Wang , Yiqiao Jin , Changyu Chen , Xinhao Zhang , Kunpeng Liu

The recent progress in Vision-Language Models (VLMs) has broadened the scope of multimodal applications. However, evaluations often remain limited to functional tasks, neglecting abstract dimensions such as personality traits and human…

Computation and Language · Computer Science 2025-06-04 Jingxuan Li , Yuning Yang , Shengqi Yang , Linfan Zhang , Ying Nian Wu

Big models, exemplified by Large Language Models (LLMs), are models typically pre-trained on massive data and comprised of enormous parameters, which not only obtain significantly improved performance across diverse tasks but also present…

Artificial Intelligence · Computer Science 2023-09-06 Jing Yao , Xiaoyuan Yi , Xiting Wang , Jindong Wang , Xing Xie

Human feedback is increasingly used to steer the behaviours of Large Language Models (LLMs). However, it is unclear how to collect and incorporate feedback in a way that is efficient, effective and unbiased, especially for highly subjective…

Computation and Language · Computer Science 2023-10-12 Hannah Rose Kirk , Andrew M. Bean , Bertie Vidgen , Paul Röttger , Scott A. Hale

Artificial intelligence (AI) is advancing at a pace that raises urgent questions about how to align machine decision-making with human moral values. This working paper investigates how leading AI systems prioritize moral outcomes and what…

Artificial Intelligence · Computer Science 2025-09-15 Eoin O'Doherty , Nicole Weinrauch , Andrew Talone , Uri Klempner , Xiaoyuan Yi , Xing Xie , Yi Zeng

This study examines the understudied role of algorithmic evaluation of human judgment in hybrid decision-making systems, a critical gap in management research. While extant literature focuses on human reluctance to follow algorithmic…

Human-Computer Interaction · Computer Science 2025-04-22 Yuanjun Feng , Vivek Chodhary , Yash Raj Shrestha

The launch of ChatGPT in November 2022 marked the beginning of a new era in AI, the availability of generative AI tools for everyone to use. ChatGPT and other similar chatbots boast a wide range of capabilities from answering student…

Computation and Language · Computer Science 2024-09-04 Yanchen Wang , Lisa Singh

This paper surveys evaluation techniques to enhance the trustworthiness and understanding of Large Language Models (LLMs). As reliance on LLMs grows, ensuring their reliability, fairness, and transparency is crucial. We explore algorithmic…

Computation and Language · Computer Science 2024-06-05 Nik Bear Brown

As LLMs occupy an increasingly important role in society, they are more and more confronted with questions that require them not only to draw on their general knowledge but also to align with certain human value systems. Therefore, studying…

Computation and Language · Computer Science 2025-10-31 Mehar Bhatia , Shravan Nayak , Gaurav Kamath , Marius Mosbach , Karolina Stańczak , Vered Shwartz , Siva Reddy
‹ Prev 1 4 5 6 7 8 10 Next ›