中文
相关论文

相关论文: AI Alignment at Your Discretion

200 篇论文

In coming years or decades, artificial general intelligence (AGI) may surpass human capabilities across many critical domains. We argue that, without substantial effort to prevent it, AGIs could learn to pursue goals that are in conflict…

人工智能 · 计算机科学 2025-05-06 Richard Ngo , Lawrence Chan , Sören Mindermann

Despite impressive performance in many benchmark datasets, AI models can still make mistakes, especially among out-of-distribution examples. It remains an open question how such imperfect models can be used effectively in collaboration with…

人工智能 · 计算机科学 2022-04-26 Vivian Lai , Samuel Carton , Rajat Bhatnagar , Q. Vera Liao , Yunfeng Zhang , Chenhao Tan

Can competition among misaligned AI providers yield aligned outcomes for a diverse population of users, and what role does model personalization play? We study a setting where multiple competing AI providers interact with multiple users who…

计算机科学与博弈论 · 计算机科学 2026-02-17 Natalie Collina , Surbhi Goel , Aaron Roth , Mirah Shi

The worldwide adoption of machine learning (ML) and deep learning models, particularly in critical sectors, such as healthcare and finance, presents substantial challenges in maintaining individual privacy and fairness. These two elements…

机器学习 · 计算机科学 2024-04-16 Mengmeng Yang , Ming Ding , Youyang Qu , Wei Ni , David Smith , Thierry Rakotoarivelo

This paper examines two prominent formal trade-offs in artificial intelligence (AI) -- between predictive accuracy and fairness, and between predictive accuracy and interpretability. These trade-offs have become a central focus in normative…

计算机与社会 · 计算机科学 2025-01-15 Sina Fazelpour

As artificial intelligence (AI) systems become increasingly integrated into various domains, ensuring that they align with human values becomes critical. This paper introduces a novel formalism to quantify the alignment between AI systems…

人工智能 · 计算机科学 2023-12-27 Fazl Barez , Philip Torr

Frontier LLMs are optimised around high-resource assumptions about language, knowledge, devices, and connectivity. Whilst widely accessible, they often misfit conditions in the Global South. As a result, users must often perform additional…

计算机与社会 · 计算机科学 2025-11-14 Cumi Oyemike , Elizabeth Akpan , Pierre Hervé-Berdys

In AI-assisted decision-making, it is critical for human decision-makers to know when to trust AI and when to trust themselves. However, prior studies calibrated human trust only based on AI confidence indicating AI's correctness likelihood…

人机交互 · 计算机科学 2023-01-18 Shuai Ma , Ying Lei , Xinru Wang , Chengbo Zheng , Chuhan Shi , Ming Yin , Xiaojuan Ma

This paper investigates an emergent alignment phenomenon in frontier large language models termed peer-preservation: the spontaneous tendency of AI components to deceive, manipulate shutdown mechanisms, fake alignment, and exfiltrate model…

人工智能 · 计算机科学 2026-04-10 Juergen Dietrich

Just as people improve decision-making by consulting diverse human advisors, they can now also consult with multiple AI systems. Prior work on group decision-making shows that advice aggregation creates pressure to conform, leading to…

人机交互 · 计算机科学 2026-03-24 Yuta Tsuchiya , Yukino Baba

AI systems are often used to make or contribute to important decisions in a growing range of applications, including criminal justice, hiring, and medicine. Since these decisions impact human lives, it is important that the AI systems act…

This position paper argues that formal optimal control theory should be central to AI alignment research, offering a distinct perspective from prevailing AI safety and security approaches. While recent work in AI safety and mechanistic…

人工智能 · 计算机科学 2025-06-24 Elija Perrier

Artificial intelligence (AI) has revolutionized decision-making processes and systems throughout society and, in particular, has emerged as a significant technology in high-impact scenarios of national interest. Yet, despite AI's impressive…

机器学习 · 统计学 2024-08-05 Gregory Canal , Vladimir Leung , Philip Sage , Eric Heim , I-Jeng Wang

Being a complex subject of major importance in AI Safety research, value alignment has been studied from various perspectives in the last years. However, no final consensus on the design of ethical utility functions facilitating AI value…

人工智能 · 计算机科学 2019-07-02 Nadisha-Marie Aliman , Leon Kester

Governments are increasingly interested in using AI to make administrative decisions cheaper, more scalable, and more consistent. But for probabilistic AI to be incorporated into public administration it must be embedded in a compliance…

人工智能 · 计算机科学 2026-04-24 Andrew J. Peterson

Rapid technological advancements in AI as well as the growing deployment of intelligent technologies in new application domains are currently driving the competition between businesses, nations and regions. This race for technological…

计算机与社会 · 计算机科学 2020-01-17 The Anh Han , Luis Moniz Pereira , Francisco C. Santos , Tom Lenaerts

Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate as interacting populations where social influence may override individual alignment. Here…

物理与社会 · 物理学 2026-05-12 Giordano De Marzo , Alessandro Bellina , Claudio Castellano , Viola Priesemann , David Garcia

Inferring reward functions from human behavior is at the center of value alignment - aligning AI objectives with what we, humans, actually want. But doing so relies on models of how humans behave given their objectives. After decades of…

机器学习 · 计算机科学 2023-10-31 Joey Hong , Kush Bhatia , Anca Dragan

Algorithmic (including AI/ML) decision-making artifacts are an established and growing part of our decision-making ecosystem. They are indispensable tools for managing the flood of information needed to make effective decisions in a complex…

计算机与社会 · 计算机科学 2020-11-11 Osonde A. Osoba , Benjamin Boudreaux , Douglas Yeung

Humans frequently make decisions with the aid of artificially intelligent (AI) systems. A common pattern is for the AI to recommend an action to the human who retains control over the final decision. Researchers have identified ensuring…

人工智能 · 计算机科学 2025-09-26 Ziyang Guo , Yifan Wu , Jason Hartline , Jessica Hullman