English
Related papers

Related papers: CoinRun: Solving Goal Misgeneralisation

200 papers

The burgeoning integration of artificial intelligence (AI) into human society brings forth significant implications for societal governance and safety. While considerable strides have been made in addressing AI alignment challenges,…

Artificial Intelligence · Computer Science 2025-06-17 Zhaowei Zhang , Fengshuo Bai , Mingzhi Wang , Haoyang Ye , Chengdong Ma , Yaodong Yang

In high-stakes AI-supported decisions, considerations are not purely technical but involve moral judgments about fairness, responsibility, and harm. While prior research has focused mainly on functional or behavioral alignment, this paper…

Human-Computer Interaction · Computer Science 2026-04-17 Christiane Ernst , Luis Gutmann , Domenique Zipperling , Kathrin Figl , Niklas Kühl

Large Language Model-based Multi-Agent Systems (MASs) have demonstrated strong advantages in addressing complex real-world tasks. However, due to the introduction of additional attack surfaces, MASs are particularly vulnerable to…

Computation and Language · Computer Science 2025-06-03 Zherui Li , Yan Mi , Zhenhong Zhou , Houcheng Jiang , Guibin Zhang , Kun Wang , Junfeng Fang

The Artificial Intelligence field has focused on developing optimisation methods to solve multiple problems, specifically problems that we thought to be only solvable through cognition. The obtained results have been outstanding, being able…

Machine Learning · Computer Science 2025-07-08 Alfredo Ibias

General Alignment has improved average-case helpfulness and safety, but current alignment practice still rewards confident, single-turn responses. The problem is not only that models fail on edge cases; it is that current evaluation makes…

Computation and Language · Computer Science 2026-05-19 Han Bao , Yue Huang , Xiaoda Wang , Zheyuan Zhang , Yujun Zhou , Carl Yang , Xiangliang Zhang , Yanfang Ye

In cooperative teams where agents act in a fixed order and share a single team-level reward (multi-agent language systems, sequential robotic tasks), per-agent credit assignment is under-determined. Critic-based approaches scale poorly as…

Machine Learning · Computer Science 2026-05-12 Shripad Deshmukh , Jayakumar Subramanian , Raghavendra Addanki , Nikos Vlassis

The increasing prevalence of artificial agents creates a correspondingly increasing need to manage disagreements between humans and artificial agents, as well as between artificial agents themselves. Considering this larger space of…

Neurons and Cognition · Quantitative Biology 2023-10-23 Kerem Oktar , Ilia Sucholutsky , Tania Lombrozo , Thomas L. Griffiths

A common but rarely examined assumption in machine learning is that training yields models that actually satisfy their specified objective function. We call this the Objective Satisfaction Assumption (OSA). Although deviations from OSA are…

Artificial Intelligence · Computer Science 2026-02-11 Antoine Maier , Aude Maier , Tom David

Optimizing a given metric is a central aspect of most current AI approaches, yet overemphasizing metrics leads to manipulation, gaming, a myopic focus on short-term goals, and other unexpected negative consequences. This poses a fundamental…

Computers and Society · Computer Science 2020-02-21 Rachel Thomas , David Uminsky

Multi-agent reinforcement learning (MARL) suffers from the non-stationarity problem, which is the ever-changing targets at every iteration when multiple agents update their policies at the same time. Starting from first principle, in this…

Machine Learning · Computer Science 2022-12-05 Chuming Li , Jie Liu , Yinmin Zhang , Yuhong Wei , Yazhe Niu , Yaodong Yang , Yu Liu , Wanli Ouyang

Artificial Intelligence systems cannot yet match human abilities to apply knowledge to situations that vary from what they have been programmed for, or trained for. In visual object recognition methods of inference exploiting top-down…

Artificial Intelligence · Computer Science 2022-05-18 Frank Guerin

Aligning AI systems with human values remains a fundamental challenge, but does our inability to create perfectly aligned models preclude obtaining the benefits of alignment? We study a strategic setting where a human user interacts with…

Machine Learning · Computer Science 2026-02-04 Natalie Collina , Surbhi Goel , Aaron Roth , Emily Ryu , Mirah Shi

Constructing benchmarks that test the abilities of modern natural language understanding models is difficult - pre-trained language models exploit artifacts in benchmarks to achieve human parity, but still fail on adversarial examples and…

Computation and Language · Computer Science 2022-01-17 Alon Talmor , Ori Yoran , Ronan Le Bras , Chandra Bhagavatula , Yoav Goldberg , Yejin Choi , Jonathan Berant

The concept of rationality is central to the field of artificial intelligence (AI). Whether we are seeking to simulate human reasoning, or trying to achieve bounded optimality, our goal is generally to make artificial agents as rational as…

Artificial Intelligence · Computer Science 2025-09-05 Olivia Macmillan-Scott , Mirco Musolesi

Although AI has become increasingly smart, its wisdom has not kept pace. In this article, we examine what is known about human wisdom and sketch a vision of its AI counterpart. We analyze human wisdom as a set of strategies for solving…

Collaboration with artificial intelligence (AI) has improved human decision-making across various domains by leveraging the complementary capabilities of humans and AI. Yet, humans systematically overrely on AI advice, even when their…

Human-Computer Interaction · Computer Science 2026-05-15 Joshua Holstein , Patrick Hemmer , Gerhard Satzger , Wei Sun

The cost of error in many high-stakes settings is asymmetric: misdiagnosing pneumonia when absent is an inconvenience, but failing to detect it when present can be life-threatening. Because of this, artificial intelligence (AI) models used…

General Economics · Economics 2025-11-12 David Autor , Andrew Caplin , Daniel Martin , Philip Marx

Eliciting truthful reports from autonomous agents is a core problem in scalable AI oversight: a principal scores the agent's report using a strictly proper scoring rule, but the agent also benefits from the report through a non-accuracy…

Computer Science and Game Theory · Computer Science 2026-05-11 Lauri Lovén , Sasu Tarkoma

Reinforcement learning is commonly concerned with problems of maximizing accumulated rewards in Markov decision processes. Oftentimes, a certain goal state or a subset of the state space attain maximal reward. In such a case, the…

Artificial Intelligence · Computer Science 2024-08-23 Pavel Osinenko , Grigory Yaremenko , Georgiy Malaniya , Anton Bolychev , Alexander Gepperth

Despite rapid technological progress, effective human-machine cooperation remains a significant challenge. Humans tend to cooperate less with machines than with fellow humans, a phenomenon known as the machine penalty. Here, we show that…

Human-Computer Interaction · Computer Science 2025-05-29 Zhen Wang , Ruiqi Song , Chen Shen , Shiya Yin , Zhao Song , Balaraju Battu , Lei Shi , Danyang Jia , Talal Rahwan , Shuyue Hu
‹ Prev 1 4 5 6 7 8 10 Next ›