English
Related papers

Related papers: Towards Measuring Goal-Directedness in AI Systems

200 papers

The field of AI alignment is concerned with AI systems that pursue unintended goals. One commonly studied mechanism by which an unintended goal might arise is specification gaming, in which the designer-provided specification is flawed in a…

Machine Learning · Computer Science 2022-11-03 Rohin Shah , Vikrant Varma , Ramana Kumar , Mary Phuong , Victoria Krakovna , Jonathan Uesato , Zac Kenton

To what extent do LLMs use their capabilities towards their given goal? We take this as a measure of their goal-directedness. We evaluate goal-directedness on tasks that require information gathering, cognitive effort, and plan execution,…

Artificial Intelligence · Computer Science 2025-04-17 Tom Everitt , Cristina Garbacea , Alexis Bellot , Jonathan Richens , Henry Papadatos , Siméon Campos , Rohin Shah

Our ability to predict the behavior of complex agents turns on the attribution of goals. Probing for goal-directed behavior comes in two flavors: Behavioral and mechanistic. The former proposes that goal-directedness can be estimated…

Multiagent Systems · Computer Science 2025-08-20 Nina Rajcic , Anders Søgaard

As artificial intelligence (AI) systems become increasingly integrated into various domains, ensuring that they align with human values becomes critical. This paper introduces a novel formalism to quantify the alignment between AI systems…

Artificial Intelligence · Computer Science 2023-12-27 Fazl Barez , Philip Torr

Goal-conditioned policies are used in order to break down complex reinforcement learning (RL) problems by using subgoals, which can be defined either in state space or in a latent feature space. This can increase the efficiency of learning…

Machine Learning · Computer Science 2020-06-04 Srinivas Venkattaramanujam , Eric Crawford , Thang Doan , Doina Precup

Recent applications of autonomous agents and robots, such as self-driving cars, scenario-based trainers, exploration robots, and service robots have brought attention to crucial trust-related challenges associated with the current…

Robotics · Computer Science 2022-09-26 Fatai Sado , Chu Kiong Loo , Wei Shiung Liew , Matthias Kerzel , Stefan Wermter

AI policy sets boundaries on acceptable behavior for AI models, but this is challenging in the context of large language models (LLMs): how do you ensure coverage over a vast behavior space? We introduce policy maps, an approach to AI…

Human-Computer Interaction · Computer Science 2025-08-04 Michelle S. Lam , Fred Hohman , Dominik Moritz , Jeffrey P. Bigham , Kenneth Holstein , Mary Beth Kery

Reinforcement learning is a promising framework for solving control problems, but its use in practical situations is hampered by the fact that reward functions are often difficult to engineer. Specifying goals and tasks for autonomous…

Machine Learning · Computer Science 2019-02-22 Justin Fu , Anoop Korattikara , Sergey Levine , Sergio Guadarrama

Goal-conditioned policy learning for robotic manipulation presents significant challenges in maintaining performance across diverse objectives and environments. We introduce Hyper-GoalNet, a framework that generates task-specific policy…

Robotics · Computer Science 2025-12-02 Pei Zhou , Wanting Yao , Qian Luo , Xunzhe Zhou , Yanchao Yang

This research explores how human-defined goals influence the behavior of Large Language Models (LLMs) through purpose-conditioned cognition. Using financial prediction tasks, we show that revealing the downstream use (e.g., predicting stock…

General Finance · Quantitative Finance 2026-05-07 Sean Cao , Wei Jiang , Hui Xu

Given that Artificial Intelligence (AI) increasingly permeates our lives, it is critical that we systematically align AI objectives with the goals and values of humans. The human-AI alignment problem stems from the impracticality of…

Computers and Society · Computer Science 2022-07-05 John Nay , James Daily

To understand the goals and goal representations of AI systems, we carefully study a pretrained reinforcement learning policy that solves mazes by navigating to a range of target squares. We find this network pursues multiple…

Artificial Intelligence · Computer Science 2023-10-13 Ulisse Mini , Peli Grietzer , Mrinank Sharma , Austin Meek , Monte MacDiarmid , Alexander Matt Turner

Manipulation is a common concern in many domains, such as social media, advertising, and chatbots. As AI systems mediate more of our interactions with the world, it is important to understand the degree to which AI systems might manipulate…

Computers and Society · Computer Science 2023-10-31 Micah Carroll , Alan Chan , Henry Ashton , David Krueger

Robotic systems are more present in our society everyday. In human-robot environments, it is crucial that end-users may correctly understand their robotic team-partners, in order to collaboratively complete a task. To increase action…

Artificial Intelligence · Computer Science 2021-09-03 Francisco Cruz , Richard Dazeley , Peter Vamplew , Ithan Moreira

As artificial intelligence (AI) becomes more powerful and widespread, the AI alignment problem - how to ensure that AI systems pursue the goals that we want them to pursue - has garnered growing attention. This article distinguishes two…

Computers and Society · Computer Science 2022-05-10 Anton Korinek , Avital Balwit

Frontier AI systems perform best in settings with clear, stable, and verifiable objectives, such as code generation, mathematical reasoning, games, and unit-test-driven tasks. They remain less reliable in open-ended settings, including…

Artificial Intelligence · Computer Science 2026-05-06 Jie Zhou , Qin Chen , Liang He

Helping people identify and pursue personally meaningful career goals at scale remains a key challenge in applied psychology. Career coaching can improve goal quality and attainment, but its cost and limited availability restrict access.…

Human-Computer Interaction · Computer Science 2026-03-19 Michel Schimpf , Julian Voigt , Thomas Bohné

Unsupervised pretraining has driven empirical advances in goal-conditioned reinforcement learning (GCRL), but its theoretical foundations remain poorly understood. In particular, an influential class of methods, mutual information skill…

Machine Learning · Computer Science 2026-05-08 Alireza Modirshanechi , Benjamin Eysenbach , Peter Dayan , Eric Schulz

Goal-conditioned reinforcement learning endows an agent with a large variety of skills, but it often struggles to solve tasks that require more temporally extended reasoning. In this work, we propose to incorporate imagined subgoals into…

Machine Learning · Computer Science 2021-07-02 Elliot Chane-Sane , Cordelia Schmid , Ivan Laptev

As language models (LMs) are increasingly deployed as autonomous agents, their robust adherence to human-assigned objectives becomes crucial for safe operation. When these agents operate independently for extended periods without human…

Artificial Intelligence · Computer Science 2025-05-06 Rauno Arike , Elizabeth Donoway , Henning Bartsch , Marius Hobbhahn
‹ Prev 1 2 3 10 Next ›