中文
相关论文

相关论文: The Alignment Trap: Complexity Barriers

200 篇论文

The AI alignment problem, which focusses on ensuring that artificial intelligence (AI), including AGI and ASI, systems act according to human values, presents profound challenges. With the progression from narrow AI to Artificial General…

人工智能 · 计算机科学 2025-07-25 Alberto Hernández-Espinosa , Felipe S. Abrahão , Olaf Witkowski , Hector Zenil

AI safety is still largely framed as alignment: training models to follow human preferences, safety policies, and normative constraints. That framing has improved the behavior of modern language models, but aligned behavior does not by…

人工智能 · 计算机科学 2026-05-27 Yige Li , Yunhao Feng , Jun Sun

With the increased deployment of large language models (LLMs), one concern is their potential misuse for generating harmful content. Our work studies the alignment challenge, with a focus on filters to prevent the generation of unsafe…

人工智能 · 计算机科学 2025-07-11 Sarah Ball , Greg Gluch , Shafi Goldwasser , Frauke Kreuter , Omer Reingold , Guy N. Rothblum

Ensuring that artificial intelligence (AI) systems satisfy formal safety and policy constraints is a central challenge in safety-critical domains. While limitations of verification are often attributed to combinatorial complexity and model…

人工智能 · 计算机科学 2026-04-07 Munawar Hasan

This paper looks at philosophical questions that arise in the context of AI alignment. It defends three propositions. First, normative and technical aspects of the AI alignment problem are interrelated, creating space for productive…

计算机与社会 · 计算机科学 2020-10-07 Iason Gabriel

A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue that, even when research agents are not scheming to…

人工智能 · 计算机科学 2026-05-18 Aleksandr Bowkis , Marie Davidsen Buhl , Jacob Pfau , Geoffrey Irving

If AI systems match or exceed human capabilities on a wide range of tasks, it may become difficult for humans to efficiently judge their actions -- making it hard to use human feedback to steer them towards desirable traits. One proposed…

人工智能 · 计算机科学 2025-05-26 Marie Davidsen Buhl , Jacob Pfau , Benjamin Hilton , Geoffrey Irving

Fairness in AI has garnered quite some attention in research, and increasingly also in society. The so-called "Impossibility Theorem" has been one of the more striking research results with both theoretical and practical consequences, as it…

计算机与社会 · 计算机科学 2023-04-14 MaryBeth Defrance , Tijl De Bie

Large Language Models (LLMs) have benefited enormously from scaling, yet these gains are bounded by five fundamental limitations: (1) hallucination, (2) context compression, (3) reasoning degradation, (4) retrieval fragility, and (5)…

If two experts disagree on a test, we may conclude both cannot be 100 per cent correct. But if they completely agree, no possible evaluation can be excluded. This asymmetry in the utility of agreements versus disagreements is explored here…

人工智能 · 计算机科学 2025-10-02 Andrés Corrada-Emmanuel

While much research in artificial intelligence (AI) has focused on scaling capabilities, the accelerating pace of development makes countervailing work on producing harmless, "aligned" systems increasingly urgent. Yet research on alignment…

人工智能 · 计算机科学 2025-12-12 Dani Roytburg , Beck Miller

We sketch how developers of frontier AI systems could construct a structured rationale -- a 'safety case' -- that an AI system is unlikely to cause catastrophic outcomes through scheming. Scheming is a potential threat model where AI…

We study reinforcement learning from human feedback under misspecification. Sometimes human feedback is systematically wrong on certain types of inputs, like a broken compass that points the wrong way in specific regions. We prove that when…

人工智能 · 计算机科学 2025-09-16 Madhava Gaikwad

We consider that existing approaches to AI "safety" and "alignment" may not be using the most effective tools, teams, or approaches. We suggest that an alternative and better approach to the problem may be to treat alignment as a social…

计算机与社会 · 计算机科学 2024-07-16 David Rostcheck , Lara Scheibling

AI Safety has become a vital front-line concern of many scientists within and outside the AI community. There are many immediate and long term anticipated risks that range from existential risk to human existence to deep fakes and bias in…

人工智能 · 计算机科学 2024-10-15 Simon Kasif

In AI alignment, extensive latitude must be granted to annotators, either human or algorithmic, to judge which model outputs are `better' or `safer.' We refer to this latitude as alignment discretion. Such discretion remains largely…

Can AI agents deal with hard choices -- cases where options are incommensurable because multiple objectives are pursued simultaneously? Adopting a technologically engaged approach distinct from existing philosophical literature, I submit…

人工智能 · 计算机科学 2026-04-21 Kangyu Wang

The critical inquiry pervading the realm of Philosophy, and perhaps extending its influence across all Humanities disciplines, revolves around the intricacies of morality and normativity. Surprisingly, in recent years, this thematic thread…

人工智能 · 计算机科学 2024-06-19 Nicholas Kluge Corrêa

Value alignment problems arise in scenarios where the specified objectives of an AI agent don't match the true underlying objective of its users. The problem has been widely argued to be one of the central safety problems in AI.…

人工智能 · 计算机科学 2023-02-10 Malek Mechergui , Sarath Sreedharan

Current large language models (LLMs) excel in verifiable domains where outputs can be checked before action but prove less reliable for high-stakes strategic decisions with uncertain outcomes. This gap, driven by mutually reinforcing…

人工智能 · 计算机科学 2025-11-12 Alejandro R. Jadad