中文
相关论文

相关论文: Goal Misgeneralization: Why Correct Specifications…

200 篇论文

Developers often struggle to specify correct training labels and rewards. Perhaps they don't need to. We propose recontextualization, which reduces how often language models "game" training signals, performing misbehaviors those signals…

Artificial Intelligence (AI) systems are increasingly used in high-stakes domains of our life, increasing the need to explain these decisions and to make sure that they are aligned with how we want the decision to be made. The field of…

人工智能 · 计算机科学 2023-06-28 Sofie Goethals , David Martens , Theodoros Evgeniou

Recent advances in Machine Learning (ML) and Artificial Intelligence (AI) follow a familiar structure: A firm releases a large, pretrained model. It is designed to be adapted and tweaked by other entities to perform particular,…

计算机科学与博弈论 · 计算机科学 2025-01-03 Benjamin Laufer , Jon Kleinberg , Hoda Heidari

Advanced AI systems sometimes act in ways that differ from human intent. To gather clear, reproducible examples, we ran the Misalignment Bounty: a crowdsourced project that collected cases of agents pursuing unintended or unsafe goals. The…

人工智能 · 计算机科学 2025-11-06 Rustem Turtayev , Natalia Fedorova , Oleg Serikov , Sergey Koldyba , Lev Avagyan , Dmitrii Volkov

Finding general evaluation metrics for unsupervised representation learning techniques is a challenging open research question, which recently has become more and more necessary due to the increasing interest in unsupervised methods. Even…

计算机视觉与模式识别 · 计算机科学 2022-03-02 Bonifaz Stuhr , Jürgen Brauer

Providing well-calibrated AI confidence can help promote users' appropriate trust in and reliance on AI, which are essential for AI-assisted decision-making. However, calibrating AI confidence -- providing confidence score that accurately…

人工智能 · 计算机科学 2025-09-30 Jingshu Li , Yitian Yang , Renwen Zhang , Q. Vera Liao , Tianqi Song , Zhengtao Xu , Yi-chieh Lee

Recently, there has been great progress in the ability of artificial intelligence (AI) algorithms to classify dermatological conditions from clinical photographs. However, little is known about the robustness of these algorithms in…

The Artificial Intelligence field has focused on developing optimisation methods to solve multiple problems, specifically problems that we thought to be only solvable through cognition. The obtained results have been outstanding, being able…

机器学习 · 计算机科学 2025-07-08 Alfredo Ibias

Goals for planning problems are typically conceived of as subsets of the state space. However, for many practical planning problems in robotics, we expect the robot to predict goals, e.g. from noisy sensors or by generalizing learned models…

机器人学 · 计算机科学 2025-07-01 Adam Conkey , Tucker Hermans

The value alignment problem for artificial intelligence (AI) is often framed as a purely technical or normative challenge, sometimes focused on hypothetical future systems. I argue that the problem is better understood as a structural…

计算机与社会 · 计算机科学 2026-04-23 Travis LaCroix

We conduct an incentivized laboratory experiment to study people's perception of generative artificial intelligence (GenAI) alignment in the context of economic decision-making. Using a panel of economic problems spanning the domains of…

理论经济学 · 经济学 2026-04-03 Kevin He , Ran Shorrer , Mengjia Xia

Although Generative Adversarial Networks achieve state-of-the-art results on a variety of generative tasks, they are regarded as highly unstable and prone to miss modes. We argue that these bad behaviors of GANs are due to the very…

机器学习 · 计算机科学 2017-03-03 Tong Che , Yanran Li , Athul Paul Jacob , Yoshua Bengio , Wenjie Li

Given that AI systems are set to play a pivotal role in future decision-making processes, their trustworthiness and reliability are of critical concern. Due to their scale and complexity, modern AI systems resist direct interpretation, and…

人工智能 · 计算机科学 2025-01-03 Binxia Xu , Antonis Bikakis , Daniel Onah , Andreas Vlachidis , Luke Dickens

The widespread adoption of large language and vision models in real-world applications has made urgent the need to address hallucinations -- instances where models produce incorrect or nonsensical outputs. These errors can propagate…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Zhengyi Ho , Siyuan Liang , Dacheng Tao

AI-related incidents are becoming increasingly frequent and severe, ranging from safety failures to misuse by malicious actors. In such complex situations, identifying which elements caused an adverse outcome, the problem of cause…

人工智能 · 计算机科学 2026-03-17 Maria Victoria Carro , David Lagnado

Misaligned research objectives have considerably hindered progress in adversarial robustness research over the past decade. For instance, an extensive focus on optimizing target metrics, while neglecting rigorous standardized evaluation,…

AI models that predict the future behavior of a system (a.k.a. predictive AI models) are central to intelligent decision-making. However, decision-making using predictive AI models often results in suboptimal performance. This is primarily…

人工智能 · 计算机科学 2025-01-13 Akhil S Anand , Shambhuraj Sawant , Dirk Reinhardt , Sebastien Gros

In this work, we present and analyze reported failures of artificially intelligent systems and extrapolate our analysis to future AIs. We suggest that both the frequency and the seriousness of future AI failures will steadily increase. AI…

人工智能 · 计算机科学 2016-10-26 Roman V. Yampolskiy , M. S. Spellchecker

An increasing number of decisions regarding the daily lives of human beings are being controlled by artificial intelligence (AI) algorithms in spheres ranging from healthcare, transportation, and education to college admissions,…

计算机与社会 · 计算机科学 2020-01-28 Dana Pessach , Erez Shmueli

As artificial intelligence scales, the concepts of alignment, agency, and autonomy have become central to AI safety, governance, and control. However, even in human contexts, these terms lack universal definitions, varying across…

计算机与社会 · 计算机科学 2025-03-11 Krti Tallam