English
Related papers

Related papers: CoinRun: Solving Goal Misgeneralisation

200 papers

Recent advances in Machine Learning (ML) and Artificial Intelligence (AI) follow a familiar structure: A firm releases a large, pretrained model. It is designed to be adapted and tweaked by other entities to perform particular,…

Computer Science and Game Theory · Computer Science 2025-01-03 Benjamin Laufer , Jon Kleinberg , Hoda Heidari

We study zero-shot generalization in reinforcement learning-optimizing a policy on a set of training tasks to perform well on a similar but unseen test task. To mitigate overfitting, previous work explored different notions of invariance to…

Machine Learning · Computer Science 2024-01-17 Ev Zisselman , Itai Lavie , Daniel Soudry , Aviv Tamar

Two key challenges within Reinforcement Learning involve improving (a) agent learning within environments with sparse extrinsic rewards and (b) the explainability of agent actions. We describe a curious subgoal focused agent to address both…

Machine Learning · Computer Science 2021-04-20 Connor van Rossum , Candice Feinberg , Adam Abu Shumays , Kyle Baxter , Benedek Bartha

For an AI's training process to successfully impart a desired goal, it is important that the AI does not attempt to resist the training. However, partially learned goals will often incentivize an AI to avoid further goal updates, as most…

Artificial Intelligence · Computer Science 2025-10-20 Rubi Hudson

In human-agent teams, openly sharing goals is often assumed to enhance planning, collaboration, and effectiveness. However, direct communication of these goals is not always feasible, requiring teammates to infer their partner's intentions…

Artificial Intelligence · Computer Science 2025-05-07 Yotam Amitai , Reuth Mirsky , Ofra Amir

Purpose: Financial service companies manage huge volumes of data which requires timely error identification and resolution. The associated tasks to resolve these errors frequently put financial analyst workforces under significant pressure…

General Finance · Quantitative Finance 2025-07-04 Chris Duckworth , Zlatko Zlatev , James Sciberras , Peter Hallett , Enrico Gerding

Intelligent physical systems as embodied cognitive systems must perform high-level reasoning while concurrently managing an underlying control architecture. The link between cognition and control must manage the problem of converting…

In this paper, the situation in which a receiver has to execute a task from a quantized version of the information source of interest is considered. The task is modeled by the minimization problem of a general goal function $f(x;g)$ for…

Signal Processing · Electrical Eng. & Systems 2022-10-03 Hang Zou , Chao Zhang , Samson Lasaulce , Lucas Saludjian , Vincent Poor

The rapid advancement of artificial intelligence (AI) systems suggests that artificial general intelligence (AGI) systems may soon arrive. Many researchers are concerned that AIs and AGIs will harm humans via intentional misuse (AI-misuse)…

Artificial Intelligence · Computer Science 2023-05-31 Catalin Mitelut , Ben Smith , Peter Vamplew

Estimating the generalization error (GE) of machine learning models is fundamental, with resampling methods being the most common approach. However, in non-standard settings, particularly those where observations are not independently and…

Goal recognition (GR) involves inferring an agent's unobserved goal from a sequence of observations. This is a critical problem in AI with diverse applications. Traditionally, GR has been addressed using 'inference to the best explanation'…

Artificial Intelligence · Computer Science 2024-09-19 Abeer Alshehri , Amal Abdulrahman , Hajar Alamri , Tim Miller , Mor Vered

The emergence of increasingly sophisticated artificial intelligence (AI) systems have sparked intense debate among researchers, policymakers, and the public due to their potential to surpass human intelligence and capabilities in all…

Theoretical Economics · Economics 2023-11-13 Mehmet S. Ismail

In Multi-Goal Reinforcement Learning, an agent learns to achieve multiple goals with a goal-conditioned policy. During learning, the agent first collects the trajectories into a replay buffer, and later these trajectories are selected…

Machine Learning · Computer Science 2020-05-26 Rui Zhao , Xudong Sun , Volker Tresp

Goal recognition is the problem of recognizing the intended goal of autonomous agents or humans by observing their behavior in an environment. Over the past years, most existing approaches to goal and plan recognition have been ignoring the…

Artificial Intelligence · Computer Science 2020-05-13 Ramon Fraga Pereira

Existing neural combinatorial optimization solvers frame solution search as imitation of optimal decisions, inherently limiting their utility to single-objective minimization and static constraints. We propose GOAL, a conditioned diffusion…

Neural and Evolutionary Computing · Computer Science 2026-05-20 Xingyu Li

We present alignment problems in current forecasting platforms, such as Good Judgment Open, CSET-Foretell or Metaculus. We classify those problems as either reward specification problems or principal-agent problems, and we propose…

Computer Science and Game Theory · Computer Science 2023-02-06 Nuño Sempere , Alex Lawsen

This paper introduces the concept of hyperpolation: a way of generalising from a limited set of data points that is a peer to the more familiar concepts of interpolation and extrapolation. Hyperpolation is the task of estimating the value…

Machine Learning · Computer Science 2024-10-15 Toby Ord

Deep neural networks excel in medical imaging but remain prone to biases, leading to fairness gaps across demographic groups. We provide the first systematic exploration of Human-AI alignment and fairness in this domain. Our results show…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Haozhe Luo , Ziyu Zhou , Zixin Shu , Aurélie Pahud de Mortanges , Robert Berke , Mauricio Reyes

Artificial intelligence (AI) is advancing exponentially and is likely to have profound impacts on human wellbeing, social equity, and environmental sustainability. Here we argue that the "alignment problem" in AI research is also an…

General Economics · Economics 2026-04-30 Daniel W. O'Neill , Stefano Vrizzi , Noemi Luna Carmeno , Felix Creutzig , Jefim Vogel

Algorithmic (including AI/ML) decision-making artifacts are an established and growing part of our decision-making ecosystem. They are indispensable tools for managing the flood of information needed to make effective decisions in a complex…

Computers and Society · Computer Science 2020-11-11 Osonde A. Osoba , Benjamin Boudreaux , Douglas Yeung