中文
相关论文

相关论文: Inverse Constitutional AI: Compressing Preferences…

200 篇论文

The trustworthiness of AI decision-making systems is increasingly important. A key feature of such systems is the ability to provide recommendations for how an individual may reverse a negative decision, a problem known as algorithmic…

人工智能 · 计算机科学 2026-05-13 Drago Plecko , Collin Wang , Elias Bareinboim

Robot policies need to adapt to human preferences and/or new environments. Human experts may have the domain knowledge required to help robots achieve this adaptation. However, existing works often require costly offline re-training on…

机器学习 · 计算机科学 2023-02-28 Vivek Myers , Erdem Bıyık , Dorsa Sadigh

This paper studies AI persuasion by distinguishing between two reasons for disagreement: attention differences, where the AI detects features the decision-maker missed, and comprehension differences, where the AI and the decision-maker…

综合经济学 · 经济学 2026-02-06 Hanzhe Li , Jin Li , Ye Luo , Xiaowei Zhang

With the growing importance of AI governance, numerous high-level frameworks and principles have been articulated by policymakers, institutions, and expert communities to guide the development and application of AI. While such frameworks…

计算机与社会 · 计算机科学 2025-06-03 Stefan Pasch

Humans rely more and more on systems with AI components. The AI community typically treats human inputs as a given and optimizes AI models only. This thinking is one-sided and it neglects the fact that humans can learn, too. In this work,…

人机交互 · 计算机科学 2020-09-22 Johannes Schneider

While a vast collection of explainable AI (XAI) algorithms have been developed in recent years, they are often criticized for significant gaps with how humans produce and consume explanations. As a result, current XAI techniques are often…

人工智能 · 计算机科学 2023-08-08 Vivian Lai , Yiming Zhang , Chacha Chen , Q. Vera Liao , Chenhao Tan

AI-enhanced personality assessments are increasingly shaping hiring decisions, using affective computing to predict traits from the Big Five (OCEAN) model. However, integrating AI into these assessments raises ethical concerns, especially…

人机交互 · 计算机科学 2025-11-24 Dena F. Mujtaba , Nihar R. Mahapatra

Learning from human preferences is important for language models to match human needs and to align with human and social values. Prior works have achieved remarkable successes by learning from human feedback to understand and follow…

机器学习 · 计算机科学 2023-10-19 Hao Liu , Carmelo Sferrazza , Pieter Abbeel

We study the problem of {\em impartial selection}, a topic that lies at the intersection of computational social choice and mechanism design. The goal is to select the most popular individual among a set of community members. The input can…

计算机科学与博弈论 · 计算机科学 2021-02-19 Ioannis Caragiannis , George Christodoulou , Nicos Protopapas

The popularity and widespread use of pruning and quantization is driven by the severe resource constraints of deploying deep neural networks to environments with strict latency, memory and energy requirements. These techniques achieve high…

机器学习 · 计算机科学 2020-12-21 Sara Hooker , Nyalleng Moorosi , Gregory Clark , Samy Bengio , Emily Denton

For summarization, human preference is critical to tame outputs of the summarizer in favor of human interests, as ground-truth summaries are scarce and ambiguous. Practical settings require dynamic exchanges between human and AI agent…

We introduce GenAI-Powered Inference (GPI), a statistical framework for both causal and predictive inference using unstructured data, including text and images. GPI leverages open-source Generative Artificial Intelligence (GenAI) models --…

机器学习 · 计算机科学 2025-09-09 Kosuke Imai , Kentaro Nakamura

How AI models should deal with political topics has been discussed, but it remains challenging and requires better governance. This paper examines the governance of large language models through individual and collective deliberation,…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Tanusree Sharma , Yujin Potter , Zachary Kilhoffer , Yun Huang , Dawn Song , Yang Wang

Most deep learning recommendation models operate as black boxes, relying on latent representations that obscure their decision process. This lack of intrinsic interpretability raises concerns in applications that require transparency and…

信息检索 · 计算机科学 2026-04-07 Jinhao Pan , Bowen Wei , Ziwei Zhu

AI-mediated communication enables users to communicate more quickly and efficiently. Various systems have been proposed such as smart reply and AI-assisted writing. Yet, the heterogeneity of the forms of inputs and architectures often…

计算与语言 · 计算机科学 2024-10-16 Benjamin Towle , Ke Zhou

Despite AI's superhuman performance in a variety of domains, humans are often unwilling to adopt AI systems. The lack of interpretability inherent in many modern AI techniques is believed to be hurting their adoption, as users may not trust…

人工智能 · 计算机科学 2021-11-17 Daehwan Ahn , Abdullah Almaatouq , Monisha Gulabani , Kartik Hosanagar

With the development of AI-Generated Content (AIGC), text-to-audio models are gaining widespread attention. However, it is challenging for these models to generate audio aligned with human preference due to the inherent information density…

声音 · 计算机科学 2024-02-02 Huan Liao , Haonan Han , Kai Yang , Tianjiao Du , Rui Yang , Zunnan Xu , Qinmei Xu , Jingquan Liu , Jiasheng Lu , Xiu Li

Explainable Artificial Intelligence (XAI) aims to make machine learning models transparent and trustworthy, yet most current approaches communicate explanations visually or through text. This paper introduces an information theoretic…

人机交互 · 计算机科学 2026-02-10 Mona Rajhans , Vishal Khawarey

As generative agents become increasingly capable, alignment of their behavior with complex human values remains a fundamental challenge. Existing approaches often simplify human intent through reduction to a scalar reward, overlooking the…

机器学习 · 计算机科学 2025-07-30 Kalyan Cherukuri , Aarav Lala

This paper presents Multi-Objective Reinforcement Learning from AI Feedback (MORLAIF), a novel approach to improving the alignment and performance of language models trained using reinforcement learning from AI feedback (RLAIF). In contrast…

机器学习 · 计算机科学 2024-06-13 Marcus Williams