中文
相关论文

相关论文: Goal Misgeneralization: Why Correct Specifications…

200 篇论文

The AI-alignment problem arises when there is a discrepancy between the goals that a human designer specifies to an AI learner and a potential catastrophic outcome that does not reflect what the human designer really wants. We argue that a…

机器学习 · 计算机科学 2020-04-10 Shai Shalev-Shwartz , Shaked Shammah , Amnon Shashua

Recent advances in deep learning have brought attention to the possibility of creating advanced, general AI systems that outperform humans across many tasks. However, if these systems pursue unintended goals, there could be catastrophic…

机器学习 · 计算机科学 2024-11-25 Dylan Xu , Juan-Pablo Rivera

Value alignment problems arise in scenarios where the specified objectives of an AI agent don't match the true underlying objective of its users. The problem has been widely argued to be one of the central safety problems in AI.…

人工智能 · 计算机科学 2023-02-10 Malek Mechergui , Sarath Sreedharan

Rapid advancements in artificial intelligence (AI) have sparked growing concerns among experts, policymakers, and world leaders regarding the potential for increasingly advanced AI systems to pose existential risks. This paper reviews the…

计算机与社会 · 计算机科学 2023-10-30 Rose Hadshar

In coming years or decades, artificial general intelligence (AGI) may surpass human capabilities across many critical domains. We argue that, without substantial effort to prevent it, AGIs could learn to pursue goals that are in conflict…

人工智能 · 计算机科学 2025-05-06 Richard Ngo , Lawrence Chan , Sören Mindermann

We study goal misgeneralization, a type of out-of-distribution generalization failure in reinforcement learning (RL). Goal misgeneralization failures occur when an RL agent retains its capabilities out-of-distribution yet pursues the wrong…

机器学习 · 计算机科学 2023-01-11 Lauro Langosco , Jack Koch , Lee Sharkey , Jacob Pfau , Laurent Orseau , David Krueger

Safe generalization in reinforcement learning requires not only that a learned policy acts capably in new situations, but also that it uses its capabilities towards the pursuit of the designer's intended goal. The latter requirement may…

Reward functions, learned or manually specified, are rarely perfect. Instead of accurately expressing human goals, these reward functions are often distorted by human beliefs about how best to achieve those goals. Specifically, these reward…

机器学习 · 计算机科学 2025-07-16 Henrik Marklund , Alex Infanger , Benjamin Van Roy

Goal misgeneralisation is a key challenge in AI alignment -- the task of getting powerful Artificial Intelligences to align their goals with human intentions and human morality. In this paper, we show how the ACE (Algorithm for Concept…

人工智能 · 计算机科学 2023-11-02 Stuart Armstrong , Alexandre Maranhão , Oliver Daniels-Koch , Patrick Leask , Rebecca Gorman

Machine learning (ML) and artificial intelligence (AI) approaches are often criticized for their inherent bias and for their lack of control, accountability, and transparency. Consequently, regulatory bodies struggle with containing this…

人工智能 · 计算机科学 2025-01-06 Benjamin Roth , Pedro Henrique Luz de Araujo , Yuxi Xia , Saskia Kaltenbrunner , Christoph Korab

As artificial intelligence (AI) becomes more powerful and widespread, the AI alignment problem - how to ensure that AI systems pursue the goals that we want them to pursue - has garnered growing attention. This article distinguishes two…

计算机与社会 · 计算机科学 2022-05-10 Anton Korinek , Avital Balwit

As AI becomes more capable, we entrust it with more general and consequential tasks. The risks from failure grow more severe with increasing task scope. It is therefore important to understand how extremely capable AI models will fail: Will…

人工智能 · 计算机科学 2026-04-13 Alexander Hägele , Aryo Pradipta Gema , Henry Sleight , Ethan Perez , Jascha Sohl-Dickstein

Bias in applications of machine learning (ML) to healthcare is usually attributed to unrepresentative or incomplete data, or to underlying health disparities. This article identifies a more pervasive source of bias that affects the clinical…

机器学习 · 计算机科学 2023-08-07 Eran Tal

A common but rarely examined assumption in machine learning is that training yields models that actually satisfy their specified objective function. We call this the Objective Satisfaction Assumption (OSA). Although deviations from OSA are…

人工智能 · 计算机科学 2026-02-11 Antoine Maier , Aude Maier , Tom David

Detecting and handling misspecified objectives, such as reward functions, has been widely recognized as one of the central challenges within the domain of Artificial Intelligence (AI) safety research. However, even with the recognition of…

人工智能 · 计算机科学 2024-11-01 Malek Mechergui , Sarath Sreedharan

Creating systems that are aligned with our goals is seen as a leading approach to create safe and beneficial AI in both leading AI companies and the academic field of AI safety. We defend the view that misaligned AGI - future, generally…

计算机与社会 · 计算机科学 2025-06-05 Max Hellrigel-Holderbaum , Leonard Dung

For an AI's training process to successfully impart a desired goal, it is important that the AI does not attempt to resist the training. However, partially learned goals will often incentivize an AI to avoid further goal updates, as most…

人工智能 · 计算机科学 2025-10-20 Rubi Hudson

Optimizing a given metric is a central aspect of most current AI approaches, yet overemphasizing metrics leads to manipulation, gaming, a myopic focus on short-term goals, and other unexpected negative consequences. This poses a fundamental…

计算机与社会 · 计算机科学 2020-02-21 Rachel Thomas , David Uminsky

This research explores how human-defined goals influence the behavior of Large Language Models (LLMs) through purpose-conditioned cognition. Using financial prediction tasks, we show that revealing the downstream use (e.g., predicting stock…

综合金融 · 定量金融 2026-05-07 Sean Cao , Wei Jiang , Hui Xu

The field of AI alignment aims to steer AI systems toward human goals, preferences, and ethical principles. Its contributions have been instrumental for improving the output quality, safety, and trustworthiness of today's AI models. This…

人工智能 · 计算机科学 2024-11-26 Robert West , Roland Aydin
‹ 上一页 1 2 3 10 下一页 ›