English
Related papers

Related papers: Inverse Constitutional AI: Compressing Preferences…

200 papers

Explainable AI (XAI) is widely used to analyze AI systems' decision-making, such as providing counterfactual explanations for recourse. When unexpected explanations occur, users may want to understand the training data properties shaping…

Machine Learning · Computer Science 2025-03-26 André Artelt , Barbara Hammer

Artificial Intelligence (AI) has been used extensively in automatic decision making in a broad variety of scenarios, ranging from credit ratings for loans to recommendations of movies. Traditional design guidelines for AI models focus…

Artificial Intelligence · Computer Science 2018-09-27 Marisa Vasconcelos , Carlos Cardonha , Bernardo Gonçalves

AI alignment, the challenge of ensuring AI systems act in accordance with human values, has emerged as a critical problem in the development of systems such as foundation models and recommender systems. Still, the current dominant approach,…

Artificial Intelligence · Computer Science 2025-03-14 Benjamin Heymann

We state the problem of inverse reinforcement learning in terms of preference elicitation, resulting in a principled (Bayesian) statistical formulation. This generalises previous work on Bayesian inverse reinforcement learning and allows us…

Machine Learning · Statistics 2011-06-30 Constantin Rothkopf , Christos Dimitrakakis

Consider the decision-making setting where agents elect a panel by expressing both positive and negative preferences. Prominently, in constitutional AI, citizens democratically select a slate of ethical preferences on which a foundation…

Computer Science and Game Theory · Computer Science 2025-03-05 Sonja Kraiczy , Georgios Papasotiropoulos , Grzegorz Pierczyński , Piotr Skowron

Current AI systems minimize risk by enforcing ideological neutrality, yet this may introduce automation bias by suppressing cognitive engagement in human decision-making. We conducted randomized trials with 2,500 participants to test…

Human-Computer Interaction · Computer Science 2025-08-21 Shiyang Lai , Junsol Kim , Nadav Kunievsky , Yujin Potter , James Evans

This research addresses a fundamental question in AI: whether large language models truly understand concepts or simply recognize patterns. The authors propose bidirectional reasoning,the ability to apply transformations in both directions…

Humans increasingly interact with Artificial intelligence(AI) systems. AI systems are optimized for objectives such as minimum computation or minimum error rate in recognizing and interpreting inputs from humans. In contrast, inputs created…

Machine Learning · Computer Science 2020-03-11 Johannes Schneider

We present NOTAI.AI, an explainable framework for machine-generated text detection that extends Fast-DetectGPT by integrating curvature-based signals with neural and stylometric features in a supervised setting. The system combines 17…

Computation and Language · Computer Science 2026-03-09 Oleksandr Marchenko Breneur , Adelaide Danilov , Aria Nourbakhsh , Salima Lamsiyah

For AI systems to be useful to humans, they must understand and act in accordance with our values and preferences. Since specifying preferences is a hard task, inverse reinforcement learning (IRL) aims to develop methods that allow for…

Artificial Intelligence · Computer Science 2026-05-12 Karim Abdel Sadek , Mark Bedaywi , Rhys Gould , Stuart Russell

Explainable artificial intelligence (XAI) has predominantly focused on generating model-centric explanations that approximate the behavior of black-box models. However, such explanations often overlook a fundamental aspect of…

Machine Learning · Computer Science 2026-04-22 Salvatore Greco , Jacek Karolczak , Roman Słowiński , Jerzy Stefanowski

Explanations of an AI's function can assist human decision-makers, but the most useful explanation depends on the decision's context, referred to as the downstream task. User studies are necessary to determine the best explanations for each…

Human-Computer Interaction · Computer Science 2024-09-20 Eura Nofshin , Esther Brown , Brian Lim , Weiwei Pan , Finale Doshi-Velez

AI systems are often used to make or contribute to important decisions in a growing range of applications, including criminal justice, hiring, and medicine. Since these decisions impact human lives, it is important that the AI systems act…

Artificial Intelligence · Computer Science 2021-03-16 Duncan C McElfresh , Lok Chan , Kenzie Doyle , Walter Sinnott-Armstrong , Vincent Conitzer , Jana Schaich Borg , John P Dickerson

Customising AI technologies to each user's preferences is fundamental to them functioning well. Unfortunately, current methods require too much user involvement and fail to capture their true preferences. In fact, to avoid the nuisance of…

Information Retrieval · Computer Science 2023-08-14 Marc Serramia , Natalia Criado , Michael Luck

Perceptual estimates exhibit a reversal in bias depending on uncertainty: they shift toward prior expectations under high stimulus noise, but away from them when sensory noise dominates. The normative framework of a Bayesian observer model…

Neurons and Cognition · Quantitative Biology 2025-10-16 Hyun-Jun Jeon , Hansol Choi , Oh-Sang Kwon

AI is powerful, but it can make choices that result in objective errors, contextually inappropriate outputs, and disliked options. We need AI-resilient interfaces that help people be resilient to the AI choices that are not right, or not…

Human-Computer Interaction · Computer Science 2024-05-15 Elena L. Glassman , Ziwei Gu , Jonathan K. Kummerfeld

Proactive AI writing assistants need to predict when users want drafting help, yet we lack empirical understanding of what drives preferences. Through a factorial vignette study with 50 participants making 750 pairwise comparisons, we find…

Computation and Language · Computer Science 2026-01-09 Vivian Lai , Zana Buçinca , Nil-Jana Akpinar , Mo Houtti , Hyeonsu B. Kang , Kevin Chian , Namjoon Suh , Alex C. Williams

Can we design artificial intelligence (AI) systems that rank our social media feeds to consider democratic values such as mitigating partisan animosity as part of their objective functions? We introduce a method for translating established,…

Human-Computer Interaction · Computer Science 2024-02-16 Chenyan Jia , Michelle S. Lam , Minh Chau Mai , Jeff Hancock , Michael S. Bernstein

When prompting a language model (LM), users often expect the model to adhere to a set of behavioral principles across diverse tasks, such as producing insightful content while avoiding harmful or biased language. Instilling such principles…

Computation and Language · Computer Science 2024-05-22 Jan-Philipp Fränken , Eric Zelikman , Rafael Rafailov , Kanishk Gandhi , Tobias Gerstenberg , Noah D. Goodman

Mainstream creativity support design prioritizes compliant AI for seamless writing interactions, but concerns over inappropriate AI reliance highlight the need for designs fostering reflection on balanced AI and non-AI resource use.…

Human-Computer Interaction · Computer Science 2026-05-19 Hua Xuan Qin , Guangzhi Zhu , Mingming Fan , Pan Hui