中文
相关论文

相关论文: Inverse Constitutional AI: Compressing Preferences…

200 篇论文

AI is increasingly used to scale collective decision-making, but far less attention has been paid to how such systems can support procedural legitimacy, particularly the conditions shaping losers' consent: whether participants who do not…

人机交互 · 计算机科学 2026-04-08 Suyash Fulay , Prerna Ravi , Emily Kubin , Shrestha Mohanty , Michiel Bakker , Deb Roy

The dominant practice of AI alignment assumes (1) that preferences are an adequate representation of human values, (2) that human rationality can be understood in terms of maximizing the satisfaction of preferences, and (3) that AI systems…

人工智能 · 计算机科学 2024-11-12 Tan Zhi-Xuan , Micah Carroll , Matija Franklin , Hal Ashton

Existing methods for controlling language models, such as RLHF and Constitutional AI, involve determining which LLM behaviors are desirable and training them into a language model. However, in many cases, it is desirable for LLMs to be…

计算与语言 · 计算机科学 2024-02-14 Louis Castricato , Nathan Lile , Suraj Anand , Hailey Schoelkopf , Siddharth Verma , Stella Biderman

We propose a novel preference alignment framework for improving spoken dialogue models on real-time conversations from user interactions. Current preference learning methods primarily focus on text-based language models, and are not…

计算与语言 · 计算机科学 2025-06-27 Anne Wu , Laurent Mazaré , Neil Zeghidour , Alexandre Défossez

In many legal processes being able to action on the concrete implication of a legal question can be valuable to automating human review or signalling certain conditions (e.g., alerts around automatic renewal). To support such tasks, we…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Adam Roegiest , Radha Chitta , Jonathan Donnelly , Maya Lash , Alexandra Vtyurina , François Longtin

As AI systems become increasingly capable and influential, ensuring their alignment with human values, preferences, and goals has become a critical research focus. Current alignment methods primarily focus on designing algorithms and loss…

计算与语言 · 计算机科学 2025-05-02 Min-Hsuan Yeh , Jeffrey Wang , Xuefeng Du , Seongheon Park , Leitian Tao , Shawn Im , Yixuan Li

Machine learning models are widely used in real-world applications. However, their complexity makes it often challenging to interpret the rationale behind their decisions. Counterfactual explanations (CEs) have emerged as a viable solution…

机器学习 · 计算机科学 2024-03-04 Muhammad Suffian , Jose M. Alonso-Moral , Alessandro Bogliolo

As AI systems become increasingly integrated into high-stakes domains, enabling users to accurately interpret model behavior is critical. While AI explanations can be provided, users often struggle to reason effectively with these…

人机交互 · 计算机科学 2025-08-27 Aniket Nuthalapati , Nicholas Hinds , Brian Y. Lim , Qianwen Wang

As AI systems increasingly shape political views, defining and evaluating AI political neutrality is an urgent problem. Here, we propose a new definition of AI political neutrality and design a large-scale user study to test it, releasing a…

计算机与社会 · 计算机科学 2026-05-29 Jonathan Stray , David Zhai Yang , Steven Luo , Miu Nicole Takagi , Serina Chang

Today, social media platforms are significant sources of news and political communication, but their role in spreading misinformation has raised significant concerns. In response, these platforms have implemented various content moderation…

计算机与社会 · 计算机科学 2026-04-21 Saeedeh Mohammadi , Taha Yasseri

Governments are increasingly interested in using AI to make administrative decisions cheaper, more scalable, and more consistent. But for probabilistic AI to be incorporated into public administration it must be embedded in a compliance…

人工智能 · 计算机科学 2026-04-24 Andrew J. Peterson

Political biases in Large Language Model (LLM)-based artificial intelligence (AI) systems, such as OpenAI's ChatGPT or Google's Gemini, have been previously reported. While several prior studies have attempted to quantify these biases using…

计算机与社会 · 计算机科学 2025-03-17 David Rozado

Why do biased predictions arise? What interventions can prevent them? We evaluate 8.2 million algorithmic predictions of math performance from $\approx$400 AI engineers, each of whom developed an algorithm under a randomly assigned…

综合经济学 · 经济学 2020-12-07 Bo Cowgill , Fabrizio Dell'Acqua , Samuel Deng , Daniel Hsu , Nakul Verma , Augustin Chaintreau

As robots and digital assistants are deployed in the real world, these agents must be able to communicate their decision-making criteria to build trust, improve human-robot teaming, and enable collaboration. While the field of explainable…

人机交互 · 计算机科学 2025-04-22 Andrew Silva , Pradyumna Tambwekar , Mariah Schrum , Matthew Gombolay

Counterfactual explanations have emerged as a prominent method in Explainable Artificial Intelligence (XAI), providing intuitive and actionable insights into Machine Learning model decisions. In contrast to other traditional feature…

Empirical human-AI alignment aims to make AI systems act in line with observed human behavior. While noble in its goals, we argue that empirical alignment can inadvertently introduce statistical biases that warrant caution. This position…

人工智能 · 计算机科学 2025-05-13 Julian Rodemann , Esteban Garces Arias , Christoph Luther , Christoph Jansen , Thomas Augustin

Can neural networks be applied in voting theory, while satisfying the need for transparency in collective decisions? We propose axiomatic deep voting: a framework to build and evaluate neural networks that aggregate preferences, using the…

人工智能 · 计算机科学 2025-08-12 Levin Hornischer , Zoi Terzopoulou

The abilities of Generative-Artificial Intelligence (AI) to produce real-time, sophisticated responses across diverse contexts has promised a huge potential in physics education, particularly in providing customized feedback. In this study,…

物理教育 · 物理学 2025-08-14 Amogh Sirnoorkar , N. Sanjay Rebello

As AI agents generate increasingly sophisticated behaviors, manually encoding human preferences to guide these agents becomes more challenging. To address this, it has been suggested that agents instead learn preferences from human choice…

机器学习 · 计算机科学 2024-12-24 Henrik Marklund , Benjamin Van Roy

Explainable AI (XAI) is an active research area to interpret a neural network's decision by ensuring transparency and trust in the task-specified learned models. Recently, perturbation-based model analysis has shown better interpretation,…

计算机视觉与模式识别 · 计算机科学 2021-02-17 Mahesh Sudhakar , Sam Sattarzadeh , Konstantinos N. Plataniotis , Jongseong Jang , Yeonjeong Jeong , Hyunwoo Kim