中文
相关论文

相关论文: CoinRun: Solving Goal Misgeneralisation

200 篇论文

Recent advances in Machine Learning (ML) and Artificial Intelligence (AI) follow a familiar structure: A firm releases a large, pretrained model. It is designed to be adapted and tweaked by other entities to perform particular,…

计算机科学与博弈论 · 计算机科学 2025-01-03 Benjamin Laufer , Jon Kleinberg , Hoda Heidari

We study zero-shot generalization in reinforcement learning-optimizing a policy on a set of training tasks to perform well on a similar but unseen test task. To mitigate overfitting, previous work explored different notions of invariance to…

机器学习 · 计算机科学 2024-01-17 Ev Zisselman , Itai Lavie , Daniel Soudry , Aviv Tamar

Two key challenges within Reinforcement Learning involve improving (a) agent learning within environments with sparse extrinsic rewards and (b) the explainability of agent actions. We describe a curious subgoal focused agent to address both…

机器学习 · 计算机科学 2021-04-20 Connor van Rossum , Candice Feinberg , Adam Abu Shumays , Kyle Baxter , Benedek Bartha

For an AI's training process to successfully impart a desired goal, it is important that the AI does not attempt to resist the training. However, partially learned goals will often incentivize an AI to avoid further goal updates, as most…

人工智能 · 计算机科学 2025-10-20 Rubi Hudson

In human-agent teams, openly sharing goals is often assumed to enhance planning, collaboration, and effectiveness. However, direct communication of these goals is not always feasible, requiring teammates to infer their partner's intentions…

人工智能 · 计算机科学 2025-05-07 Yotam Amitai , Reuth Mirsky , Ofra Amir

Purpose: Financial service companies manage huge volumes of data which requires timely error identification and resolution. The associated tasks to resolve these errors frequently put financial analyst workforces under significant pressure…

综合金融 · 定量金融 2025-07-04 Chris Duckworth , Zlatko Zlatev , James Sciberras , Peter Hallett , Enrico Gerding

Intelligent physical systems as embodied cognitive systems must perform high-level reasoning while concurrently managing an underlying control architecture. The link between cognition and control must manage the problem of converting…

In this paper, the situation in which a receiver has to execute a task from a quantized version of the information source of interest is considered. The task is modeled by the minimization problem of a general goal function $f(x;g)$ for…

信号处理 · 电气工程与系统科学 2022-10-03 Hang Zou , Chao Zhang , Samson Lasaulce , Lucas Saludjian , Vincent Poor

The rapid advancement of artificial intelligence (AI) systems suggests that artificial general intelligence (AGI) systems may soon arrive. Many researchers are concerned that AIs and AGIs will harm humans via intentional misuse (AI-misuse)…

人工智能 · 计算机科学 2023-05-31 Catalin Mitelut , Ben Smith , Peter Vamplew

Estimating the generalization error (GE) of machine learning models is fundamental, with resampling methods being the most common approach. However, in non-standard settings, particularly those where observations are not independently and…

Goal recognition (GR) involves inferring an agent's unobserved goal from a sequence of observations. This is a critical problem in AI with diverse applications. Traditionally, GR has been addressed using 'inference to the best explanation'…

人工智能 · 计算机科学 2024-09-19 Abeer Alshehri , Amal Abdulrahman , Hajar Alamri , Tim Miller , Mor Vered

The emergence of increasingly sophisticated artificial intelligence (AI) systems have sparked intense debate among researchers, policymakers, and the public due to their potential to surpass human intelligence and capabilities in all…

理论经济学 · 经济学 2023-11-13 Mehmet S. Ismail

In Multi-Goal Reinforcement Learning, an agent learns to achieve multiple goals with a goal-conditioned policy. During learning, the agent first collects the trajectories into a replay buffer, and later these trajectories are selected…

机器学习 · 计算机科学 2020-05-26 Rui Zhao , Xudong Sun , Volker Tresp

Goal recognition is the problem of recognizing the intended goal of autonomous agents or humans by observing their behavior in an environment. Over the past years, most existing approaches to goal and plan recognition have been ignoring the…

人工智能 · 计算机科学 2020-05-13 Ramon Fraga Pereira

Existing neural combinatorial optimization solvers frame solution search as imitation of optimal decisions, inherently limiting their utility to single-objective minimization and static constraints. We propose GOAL, a conditioned diffusion…

神经与进化计算 · 计算机科学 2026-05-20 Xingyu Li

We present alignment problems in current forecasting platforms, such as Good Judgment Open, CSET-Foretell or Metaculus. We classify those problems as either reward specification problems or principal-agent problems, and we propose…

计算机科学与博弈论 · 计算机科学 2023-02-06 Nuño Sempere , Alex Lawsen

This paper introduces the concept of hyperpolation: a way of generalising from a limited set of data points that is a peer to the more familiar concepts of interpolation and extrapolation. Hyperpolation is the task of estimating the value…

机器学习 · 计算机科学 2024-10-15 Toby Ord

Deep neural networks excel in medical imaging but remain prone to biases, leading to fairness gaps across demographic groups. We provide the first systematic exploration of Human-AI alignment and fairness in this domain. Our results show…

计算机视觉与模式识别 · 计算机科学 2025-05-16 Haozhe Luo , Ziyu Zhou , Zixin Shu , Aurélie Pahud de Mortanges , Robert Berke , Mauricio Reyes

Artificial intelligence (AI) is advancing exponentially and is likely to have profound impacts on human wellbeing, social equity, and environmental sustainability. Here we argue that the "alignment problem" in AI research is also an…

综合经济学 · 经济学 2026-04-30 Daniel W. O'Neill , Stefano Vrizzi , Noemi Luna Carmeno , Felix Creutzig , Jefim Vogel

Algorithmic (including AI/ML) decision-making artifacts are an established and growing part of our decision-making ecosystem. They are indispensable tools for managing the flood of information needed to make effective decisions in a complex…

计算机与社会 · 计算机科学 2020-11-11 Osonde A. Osoba , Benjamin Boudreaux , Douglas Yeung