中文
相关论文

相关论文: The AI Security Zugzwang

200 篇论文

The AI-alignment problem arises when there is a discrepancy between the goals that a human designer specifies to an AI learner and a potential catastrophic outcome that does not reflect what the human designer really wants. We argue that a…

机器学习 · 计算机科学 2020-04-10 Shai Shalev-Shwartz , Shaked Shammah , Amnon Shashua

How can we ensure that AI systems are aligned with human values and remain safe? We can study this problem through the frameworks of the AI assistance and the AI shutdown games. The AI assistance problem concerns designing an AI agent that…

人工智能 · 计算机科学 2025-12-30 Alessio Benavoli , Alessandro Facchini , Marco Zaffalon

Artificial intelligence (AI) is reshaping society, from video generation to medical diagnosis, coding agents to autonomous vehicles. Yet researchers, policymakers, and technology companies lack shared terminology for discussing AI risks.…

Advances in AI are widely understood to have implications for cybersecurity. Articles have emphasized the effect of AI on the cyber offense-defense balance, and commentators can be found arguing either that cyber will privilege attackers or…

密码学与安全 · 计算机科学 2025-08-25 Benjamin Murphy , Twm Stone

Recent breakthroughs in artificial intelligence (AI) have brought about increasingly capable systems that demonstrate remarkable abilities in reasoning, language understanding, and problem-solving. These advancements have prompted a renewed…

人工智能 · 计算机科学 2025-07-01 Xiaojian Li , Haoyuan Shi , Rongwu Xu , Wei Xu

A core challenge in the development of increasingly capable AI systems is to make them safe and reliable by ensuring their behaviour is consistent with human values. This challenge, known as the alignment problem, does not merely apply to…

机器学习 · 计算机科学 2023-11-07 Raphaël Millière

Globally, artificial intelligence (AI) implementation is growing, holding the capability to fundamentally alter organisational processes and decision making. Simultaneously, this brings a multitude of emergent risks to organisations,…

计算机与社会 · 计算机科学 2024-04-10 Finlay McGee

As AI systems are integrated into high stakes social domains, researchers now examine how to design and operate them in a safe and ethical manner. However, the criteria for identifying and diagnosing safety risks in complex social contexts…

计算机与社会 · 计算机科学 2021-06-22 Roel Dobbe , Thomas Krendl Gilbert , Yonatan Mintz

In the face of rapidly advancing AI technology, individuals will increasingly rely on AI agents to navigate life's growing complexities, raising critical concerns about maintaining both human agency and autonomy. This paper addresses a…

计算机与社会 · 计算机科学 2025-04-29 Philipp Koralus

If AI systems match or exceed human capabilities on a wide range of tasks, it may become difficult for humans to efficiently judge their actions -- making it hard to use human feedback to steer them towards desirable traits. One proposed…

人工智能 · 计算机科学 2025-05-26 Marie Davidsen Buhl , Jacob Pfau , Benjamin Hilton , Geoffrey Irving

The deployment of artificial intelligence (AI) applications has accelerated rapidly. AI enabled technologies are facing the public in many ways including infrastructure, consumer products and home applications. Because many of these…

人工智能 · 计算机科学 2024-08-01 Joanna F. DeFranco , Luke Biersmith

As AI rapidly advances, the security risks posed by AI are becoming increasingly severe, especially in critical scenarios, including those posing existential risks. If AI becomes uncontrollable, manipulated, or actively evades safety…

人工智能 · 计算机科学 2025-08-29 Donglin Wang , Weiyun Liang , Chunyuan Chen , Jing Xu , Yulong Fu

As AI systems become more capable, integrated, and widespread, understanding the associated risks becomes increasingly important. This paper maps the full spectrum of AI risks, from current harms affecting individual users to existential…

计算机与社会 · 计算机科学 2025-08-20 Markov Grey , Charbel-Raphaël Segerie

Artificial intelligence (AI) has been advancing at a fast pace and it is now poised for deployment in a wide range of applications, such as autonomous systems, medical diagnosis and natural language processing. Early adoption of AI…

机器学习 · 计算机科学 2023-09-21 Marta Kwiatkowska , Xiyue Zhang

AI safety is still largely framed as alignment: training models to follow human preferences, safety policies, and normative constraints. That framing has improved the behavior of modern language models, but aligned behavior does not by…

人工智能 · 计算机科学 2026-05-27 Yige Li , Yunhao Feng , Jun Sun

Ttraditional safety engineering is coming to a turning point moving from deterministic, non-evolving systems operating in well-defined contexts to increasingly autonomous and learning-enabled AI systems which are acting in largely…

人工智能 · 计算机科学 2022-05-13 Harald Rueß , Simon Burton

This position paper contends that modern AI research must adopt an antifragile perspective on safety -- one in which the system's capacity to guarantee long-term AI safety such as handling rare or out-of-distribution (OOD) events expands…

人工智能 · 计算机科学 2025-09-18 Ming Jin , Hyunin Lee

Humanity appears to be on course to soon develop AI systems that substantially outperform human experts in all cognitive domains and activities. We believe the default trajectory has a high likelihood of catastrophe, including human…

计算机与社会 · 计算机科学 2025-05-08 Peter Barnett , Aaron Scher

The off-switch problem is a critical challenge in AI control: if an AI system resists being switched off, it poses a significant risk. In this paper, we model the off-switch problem as a signalling game, where a human decision-maker…

机器学习 · 计算机科学 2025-04-01 Alessio Benavoli , Alessandro Facchini , Marco Zaffalon

The digital age, driven by the AI revolution, brings significant opportunities but also conceals security threats, which we refer to as cyber shadows. These threats pose risks at individual, organizational, and societal levels. This paper…

密码学与安全 · 计算机科学 2025-01-29 Marc Schmitt , Pantelis Koutroumpis