English
Related papers

Related papers: RAIL in the Wild: Operationalizing Responsible AI …

200 papers

We present an ethical decision-making framework that refines a pre-trained reinforcement learning (RL) model using a task-agnostic ethical layer. Following initial training, the RL model undergoes ethical fine-tuning, where human feedback…

Computers and Society · Computer Science 2026-05-05 Rohit K. Dubey , Damian Dailisan , Sachit Mahajan

This paper develops a prudential framework for assessing the reliability of large language models (LLMs) in reinsurance. A five-pillar architecture--governance, data lineage, assurance, resilience, and regulatory alignment--translates…

Artificial Intelligence · Computer Science 2025-11-12 Stella C. Dong

Research in Responsible AI has developed a range of principles and practices to ensure that machine learning systems are used in a manner that is ethical and aligned with human values. However, a critical yet often neglected aspect of…

Computers and Society · Computer Science 2024-08-21 Neha R. Gupta , Jessica Hullman , Hari Subramonyam

Large Language Models (LLMs) demonstrate a remarkable capacity to adopt different personas and roles; however, it remains unclear whether they can manifest behavior that adheres to a coherent, human-like value structure. In this work, we…

Artificial Intelligence · Computer Science 2026-05-29 Asaf Yehudai , Naama Rozen , Ariel Gera

As Large Language Models (LLMs) transition from static tools to autonomous agents, traditional evaluation benchmarks that measure performance on downstream tasks are becoming insufficient. These methods fail to capture the emergent social…

Artificial Intelligence · Computer Science 2025-10-03 Zarreen Reza

Conversational human-likeness plays a central role in human-AI interaction, yet it has remained difficult to define, measure, and optimize. As a result, improvements in human-like behavior are largely driven by scale or broad supervised…

Artificial Intelligence · Computer Science 2026-01-08 Masum Hasan , Junjie Zhao , Ehsan Hoque

The advancement of powerful yet opaque large language models (LLMs) necessitates a fundamental revision of the philosophical criteria used to evaluate artificial moral agents (AMAs). Pre-LLM frameworks often relied on the assumption of…

Artificial Intelligence · Computer Science 2025-07-29 Matthew E. Brophy

Work on morality in large language models (LLMs) has progressed via constitutional AI, reinforcement learning from human feedback (RLHF) and systematic benchmarking, yet it still lacks tools to connect internal moral representations to…

Human-Computer Interaction · Computer Science 2026-03-25 Gunter Bombaerts

Conversational AI systems have emerged as key enablers of human-like interactions across diverse sectors. Nevertheless, the balance between linguistic nuance and factual accuracy has proven elusive. In this paper, we first introduce…

Computation and Language · Computer Science 2024-06-18 Ahtsham Zafar , Venkatesh Balavadhani Parthasarathy , Chan Le Van , Saad Shahid , Aafaq Iqbal khan , Arsalan Shahid

The alignment between humans and machines is a critical challenge in artificial intelligence today. Reinforcement learning, which aims to maximize a reward function, is particularly vulnerable to the risks associated with poorly designed…

Artificial Intelligence · Computer Science 2025-10-29 Valentin Cuzin-Rambaud , Emilien Komlenovic , Alexandre Faure , Bruno Yun

Many sets of ethics principles for responsible AI have been proposed to allay concerns about misuse and abuse of AI/ML systems. The underlying aspects of such sets of principles include privacy, accuracy, fairness, robustness,…

Computers and Society · Computer Science 2024-09-09 Conrad Sanderson , David Douglas , Qinghua Lu

While Large Language Models (LLMs) have achieved remarkable success in cognitive and reasoning benchmarks, they exhibit a persistent deficit in anthropomorphic intelligence-the capacity to navigate complex social, emotional, and ethical…

Computation and Language · Computer Science 2025-12-29 Jiaxin Liu , Peiyi Tu , Wenyu Chen , Yihong Zhuang , Xinxia Ling , Anji Zhou , Chenxi Wang , Zhuo Han , Zhengkai Yang , Junbo Zhao , Zenan Huang , Yuanyuan Wang

Artificial intelligence (AI) offers incredible possibilities for patient care, but raises significant ethical issues, such as the potential for bias. Powerful ethical frameworks exist to minimize these issues, but are often developed for…

Computers and Society · Computer Science 2025-07-04 Ion Nemteanu , Adir Mancebo , Leslie Joe , Ryan Lopez , Patricia Lopez , Warren Woodrich Pettine

In this paper, we cover approaches to systematically govern, assess and quantify bias across the complete life cycle of machine learning models, from initial development and validation to ongoing production monitoring and guardrail…

Computation and Language · Computer Science 2025-08-07 Alok Abhishek , Lisa Erickson , Tushar Bandopadhyay

Current approaches to AI safety define red lines at the case level: specific prompts, specific outputs, specific harms. This paper argues that red lines can be set more fundamentally -- at the level of value, evidence, and source…

Artificial Intelligence · Computer Science 2026-04-14 Seulki Lee

The recent rise in popularity of large language models (LLMs) has prompted considerable concerns about their moral capabilities. Although considerable effort has been dedicated to aligning LLMs with human moral values, existing benchmarks…

Artificial Intelligence · Computer Science 2025-08-19 Alessio Galatolo , Luca Alberto Rappuoli , Katie Winkle , Meriem Beloucif

Large language models (LLMs) increasingly participate in morally sensitive decision-making, yet how they organize ethical frameworks across reasoning steps remains underexplored. We introduce \textit{moral reasoning trajectories}, sequences…

Computation and Language · Computer Science 2026-03-18 Fan Huang , Haewoon Kwak , Jisun An

Although artificial intelligence (AI) is solving real-world challenges and transforming industries, there are serious concerns about its ability to behave and make decisions in a responsible way. Many AI ethics principles and guidelines for…

Artificial Intelligence · Computer Science 2022-07-22 Qinghua Lu , Liming Zhu , Xiwei Xu , Jon Whittle , David Douglas , Conrad Sanderson

Large language models (LLMs) are rapidly being integrated into psychological research as research tools, evaluation targets, human simulators, and cognitive models. However, recent evidence reveals severe measurement unreliability:…

Human-Computer Interaction · Computer Science 2025-07-08 Zhicheng Lin

We propose LLM-Eval, a unified multi-dimensional automatic evaluation method for open-domain conversations with large language models (LLMs). Existing evaluation methods often rely on human annotations, ground-truth responses, or multiple…

Computation and Language · Computer Science 2023-05-24 Yen-Ting Lin , Yun-Nung Chen