English
Related papers

Related papers: Democratizing Reward Design for Personal and Repre…

200 papers

We propose a novel preference alignment framework for improving spoken dialogue models on real-time conversations from user interactions. Current preference learning methods primarily focus on text-based language models, and are not…

Computation and Language · Computer Science 2025-06-27 Anne Wu , Laurent Mazaré , Neil Zeghidour , Alexandre Défossez

As we consider entrusting Large Language Models (LLMs) with key societal and decision-making roles, measuring their alignment with human cognition becomes critical. This requires methods that can assess how these systems represent…

Artificial Intelligence · Computer Science 2025-10-03 Mattson Ogg , Ritwik Bose , Jamie Scharf , Christopher Ratto , Michael Wolmetz

As AI systems advance in capabilities, measuring their safety and alignment to human values is becoming paramount. A fast-growing field of AI research is devoted to developing such assessments. However, most current advances therein may be…

Computers and Society · Computer Science 2026-03-17 Max Hellrigel-Holderbaum , Edward James Young

As AI systems become more advanced and widely deployed, there will likely be increasing debate over whether AI systems could have conscious experiences, desires, or other states of potential moral significance. It is important to inform…

Machine Learning · Computer Science 2023-11-16 Ethan Perez , Robert Long

"Human-aware" has become a popular keyword used to describe a particular class of AI systems that are designed to work and interact with humans. While there exists a surprising level of consistency among the works that use the label…

Robotics · Computer Science 2024-07-16 Silvia Tulli , Stylianos Loukas Vasileiou , Sarath Sreedharan

The deployment of decision-making AI agents presents a critical challenge in maintaining alignment with human values or guidelines while operating in complex, dynamic environments. Agents trained solely to achieve their objectives may adopt…

Artificial Intelligence · Computer Science 2025-12-09 Dena Mujtaba , Brian Hu , Anthony Hoogs , Arslan Basharat

Large language models (LLMs) are increasingly being used as decision aids. However, users have diverse values and preferences that can affect their decision-making, which requires novel methods for LLM alignment and personalization.…

Computation and Language · Computer Science 2025-07-15 Bharadwaj Ravichandran , David Joy , Paul Elliott , Brian Hu , Jadie Adams , Christopher Funk , Emily Veenhuis , Anthony Hoogs , Arslan Basharat

AI alignment is a field of research that aims to develop methods to ensure that agents always behave in a manner aligned with (i.e. consistently with) the goals and values of their human operators, no matter their level of capability. This…

Artificial Intelligence · Computer Science 2025-05-26 Eli Sennesh , Maxwell Ramstead

Human computer interaction is shifting from screen-based systems to multimodal interfaces where artificial intelligence powered systems increasingly interpret user intent through speech, gesture, and gaze. Yet users rarely understand how…

Human-Computer Interaction · Computer Science 2026-05-05 Ankur Bhatt , Sven Mayer

Explainable AI Planning (XAIP) aims to develop AI agents that can effectively explain their decisions and actions to human users, fostering trust and facilitating human-AI collaboration. A key challenge in XAIP is model reconciliation,…

Artificial Intelligence · Computer Science 2024-05-30 Yinxu Tang , Stylianos Loukas Vasileiou , William Yeoh

The autonomous decision-making process, which is increasingly applied to computer systems, requires that the choices made by these systems align with human values. In this context, systems must assess how well their decisions reflect human…

Computers and Society · Computer Science 2025-12-19 Eduardo de la Cruz Fernández , Marcelo Karanik , Sascha Ossowski

Smart assistants increasingly act proactively, yet mistimed or intrusive behavior often causes users to lose trust and disable these features. Learning user preferences for proactive assistance is difficult because real-world studies are…

Human-Computer Interaction · Computer Science 2026-02-05 Ziyi Xuan , Yiwen Wu , Zhaoyang Yan , Vinod Namboodiri , Yu Yang

Personalized dialogue generation aims to leverage persona profiles and dialogue history to generate persona-relevant and consistent responses. Mainstream models typically rely on token-level language model training with persona dialogue…

Computation and Language · Computer Science 2025-11-14 Guanrong Li , Xinyu Liu , Zhen Wu , Xinyu Dai

This study investigates students' perceptions of Artificial Intelligence (AI) grading systems in an undergraduate computer science course (n = 27), focusing on a block-based programming final project. Guided by the ethical principles…

Artificial Intelligence · Computer Science 2026-02-24 Bahare Riahi , Viktoriia Storozhevykh , Veronica Catete

Are AI systems truly representing human values, or merely averaging across them? Our study suggests a concerning reality: Large Language Models (LLMs) fail to represent diverse cultural moral frameworks despite their linguistic…

Computation and Language · Computer Science 2025-08-01 Simon Münker

Algorithms and other formal models purportedly incorporating human values like fairness have grown increasingly popular in computer science. In response to sociotechnical challenges in the use of these models, designers and researchers have…

Computers and Society · Computer Science 2021-06-18 Benjamin Fish , Luke Stark

We explore the task of improving persona consistency of dialogue agents. Recent models tackling consistency often train with additional Natural Language Inference (NLI) labels or attach trained extra modules to the generative agent for…

Computation and Language · Computer Science 2020-10-07 Hyunwoo Kim , Byeongchang Kim , Gunhee Kim

Generating complex behaviors that satisfy the preferences of non-expert users is a crucial requirement for AI agents. Interactive reward learning from trajectory comparisons (a.k.a. RLHF) is one way to allow non-expert users to convey…

Artificial Intelligence · Computer Science 2023-03-01 Lin Guan , Karthik Valmeekam , Subbarao Kambhampati

We study AI alignment through the lens of law-and-economics models of deterrence and enforcement. In these models, misconduct is not treated as an external failure, but as a strategic response to incentives: an actor weighs the gain from…

Machine Learning · Computer Science 2026-05-12 Rohit Agarwal , Joshua Lin , Mark Braverman , Elad Hazan

The field of AI alignment aims to steer AI systems toward human goals, preferences, and ethical principles. Its contributions have been instrumental for improving the output quality, safety, and trustworthiness of today's AI models. This…

Artificial Intelligence · Computer Science 2024-11-26 Robert West , Roland Aydin