中文
相关论文

相关论文: CoinRun: Solving Goal Misgeneralisation

200 篇论文

The burgeoning integration of artificial intelligence (AI) into human society brings forth significant implications for societal governance and safety. While considerable strides have been made in addressing AI alignment challenges,…

人工智能 · 计算机科学 2025-06-17 Zhaowei Zhang , Fengshuo Bai , Mingzhi Wang , Haoyang Ye , Chengdong Ma , Yaodong Yang

In high-stakes AI-supported decisions, considerations are not purely technical but involve moral judgments about fairness, responsibility, and harm. While prior research has focused mainly on functional or behavioral alignment, this paper…

人机交互 · 计算机科学 2026-04-17 Christiane Ernst , Luis Gutmann , Domenique Zipperling , Kathrin Figl , Niklas Kühl

Large Language Model-based Multi-Agent Systems (MASs) have demonstrated strong advantages in addressing complex real-world tasks. However, due to the introduction of additional attack surfaces, MASs are particularly vulnerable to…

计算与语言 · 计算机科学 2025-06-03 Zherui Li , Yan Mi , Zhenhong Zhou , Houcheng Jiang , Guibin Zhang , Kun Wang , Junfeng Fang

The Artificial Intelligence field has focused on developing optimisation methods to solve multiple problems, specifically problems that we thought to be only solvable through cognition. The obtained results have been outstanding, being able…

机器学习 · 计算机科学 2025-07-08 Alfredo Ibias

General Alignment has improved average-case helpfulness and safety, but current alignment practice still rewards confident, single-turn responses. The problem is not only that models fail on edge cases; it is that current evaluation makes…

计算与语言 · 计算机科学 2026-05-19 Han Bao , Yue Huang , Xiaoda Wang , Zheyuan Zhang , Yujun Zhou , Carl Yang , Xiangliang Zhang , Yanfang Ye

In cooperative teams where agents act in a fixed order and share a single team-level reward (multi-agent language systems, sequential robotic tasks), per-agent credit assignment is under-determined. Critic-based approaches scale poorly as…

机器学习 · 计算机科学 2026-05-12 Shripad Deshmukh , Jayakumar Subramanian , Raghavendra Addanki , Nikos Vlassis

The increasing prevalence of artificial agents creates a correspondingly increasing need to manage disagreements between humans and artificial agents, as well as between artificial agents themselves. Considering this larger space of…

神经元与认知 · 定量生物学 2023-10-23 Kerem Oktar , Ilia Sucholutsky , Tania Lombrozo , Thomas L. Griffiths

A common but rarely examined assumption in machine learning is that training yields models that actually satisfy their specified objective function. We call this the Objective Satisfaction Assumption (OSA). Although deviations from OSA are…

人工智能 · 计算机科学 2026-02-11 Antoine Maier , Aude Maier , Tom David

Optimizing a given metric is a central aspect of most current AI approaches, yet overemphasizing metrics leads to manipulation, gaming, a myopic focus on short-term goals, and other unexpected negative consequences. This poses a fundamental…

计算机与社会 · 计算机科学 2020-02-21 Rachel Thomas , David Uminsky

Multi-agent reinforcement learning (MARL) suffers from the non-stationarity problem, which is the ever-changing targets at every iteration when multiple agents update their policies at the same time. Starting from first principle, in this…

机器学习 · 计算机科学 2022-12-05 Chuming Li , Jie Liu , Yinmin Zhang , Yuhong Wei , Yazhe Niu , Yaodong Yang , Yu Liu , Wanli Ouyang

Artificial Intelligence systems cannot yet match human abilities to apply knowledge to situations that vary from what they have been programmed for, or trained for. In visual object recognition methods of inference exploiting top-down…

人工智能 · 计算机科学 2022-05-18 Frank Guerin

Aligning AI systems with human values remains a fundamental challenge, but does our inability to create perfectly aligned models preclude obtaining the benefits of alignment? We study a strategic setting where a human user interacts with…

机器学习 · 计算机科学 2026-02-04 Natalie Collina , Surbhi Goel , Aaron Roth , Emily Ryu , Mirah Shi

Constructing benchmarks that test the abilities of modern natural language understanding models is difficult - pre-trained language models exploit artifacts in benchmarks to achieve human parity, but still fail on adversarial examples and…

计算与语言 · 计算机科学 2022-01-17 Alon Talmor , Ori Yoran , Ronan Le Bras , Chandra Bhagavatula , Yoav Goldberg , Yejin Choi , Jonathan Berant

The concept of rationality is central to the field of artificial intelligence (AI). Whether we are seeking to simulate human reasoning, or trying to achieve bounded optimality, our goal is generally to make artificial agents as rational as…

人工智能 · 计算机科学 2025-09-05 Olivia Macmillan-Scott , Mirco Musolesi

Although AI has become increasingly smart, its wisdom has not kept pace. In this article, we examine what is known about human wisdom and sketch a vision of its AI counterpart. We analyze human wisdom as a set of strategies for solving…

Collaboration with artificial intelligence (AI) has improved human decision-making across various domains by leveraging the complementary capabilities of humans and AI. Yet, humans systematically overrely on AI advice, even when their…

人机交互 · 计算机科学 2026-05-15 Joshua Holstein , Patrick Hemmer , Gerhard Satzger , Wei Sun

The cost of error in many high-stakes settings is asymmetric: misdiagnosing pneumonia when absent is an inconvenience, but failing to detect it when present can be life-threatening. Because of this, artificial intelligence (AI) models used…

综合经济学 · 经济学 2025-11-12 David Autor , Andrew Caplin , Daniel Martin , Philip Marx

Eliciting truthful reports from autonomous agents is a core problem in scalable AI oversight: a principal scores the agent's report using a strictly proper scoring rule, but the agent also benefits from the report through a non-accuracy…

计算机科学与博弈论 · 计算机科学 2026-05-11 Lauri Lovén , Sasu Tarkoma

Reinforcement learning is commonly concerned with problems of maximizing accumulated rewards in Markov decision processes. Oftentimes, a certain goal state or a subset of the state space attain maximal reward. In such a case, the…

人工智能 · 计算机科学 2024-08-23 Pavel Osinenko , Grigory Yaremenko , Georgiy Malaniya , Anton Bolychev , Alexander Gepperth

Despite rapid technological progress, effective human-machine cooperation remains a significant challenge. Humans tend to cooperate less with machines than with fellow humans, a phenomenon known as the machine penalty. Here, we show that…

人机交互 · 计算机科学 2025-05-29 Zhen Wang , Ruiqi Song , Chen Shen , Shiya Yin , Zhao Song , Balaraju Battu , Lei Shi , Danyang Jia , Talal Rahwan , Shuyue Hu