中文
相关论文

相关论文: Out of Control -- Why Alignment Needs Formal Contr…

200 篇论文

The increasing integration of artificial intelligence (AI) systems in various fields requires solid concepts to ensure compliance with upcoming legislation. This paper systematically examines the compliance of AI systems with relevant…

计算机与社会 · 计算机科学 2026-04-21 Julius Schöning , Niklas Kruse

Automated control monitors could play an important role in overseeing highly capable AI agents that we do not fully trust. Prior work has explored control monitoring in simplified settings, but scaling monitoring to real-world deployments…

This policy report draws on country studies from China, South Korea, Singapore, and the United Kingdom to identify effective tools and key barriers to interoperability in AI safety governance. It offers practical recommendations to support…

计算机与社会 · 计算机科学 2026-01-13 Yik Chan Chin , David A. Raho , Hag-Min Kim , Chunli Bi , James Ong , Jingbo Huang , Serge Stinckwich

Ensuring fairness in the coordination of connected and automated vehicles at intersections is essential for equitable access, social acceptance, and long-term system efficiency, yet it remains underexplored in safety-critical, real-time…

机器人学 · 计算机科学 2025-11-11 Lei Shi , Yongju Kim , Xinzhi Zhong , Wissam Kontar , Qichao Liu , Soyoung Ahn

Large Language Models (LLMs) face a fundamental safety-helpfulness trade-off due to static, one-size-fits-all safety policies that lack runtime controllabilityxf, making it difficult to tailor responses to diverse application needs. %As a…

计算与语言 · 计算机科学 2026-02-09 Jianfeng Si , Lin Sun , Weihong Lin , Xiangzheng Zhang

As artificial intelligence (AI) systems become increasingly integral to critical infrastructure and global operations, the need for a unified, trustworthy governance framework is more urgent that ever. This paper proposes a novel approach…

人工智能 · 计算机科学 2025-01-17 Vikram Kulothungan

This position paper argues that achieving meaningful scientific and societal advances with artificial intelligence (AI) requires a responsible, application-driven approach (RAD) to AI research. As AI is increasingly integrated into society,…

机器学习 · 计算机科学 2025-08-20 Sarah Hartman , Cheng Soon Ong , Julia Powles , Petra Kuhnert

Safety has become the central value around which dominant AI governance efforts are being shaped. Recently, this culminated in the publication of the International AI Safety Report, written by 96 experts of which 30 nominated by the…

计算机与社会 · 计算机科学 2025-03-10 Roel Dobbe

Artificial intelligence (AI) has been advancing at a fast pace and it is now poised for deployment in a wide range of applications, such as autonomous systems, medical diagnosis and natural language processing. Early adoption of AI…

机器学习 · 计算机科学 2023-09-21 Marta Kwiatkowska , Xiyue Zhang

With increased power and prevalence of AI systems, it is ever more critical that AI systems are designed to serve all, i.e., people with diverse values and perspectives. However, aligning models to serve pluralistic human values remains an…

Recent discussions and research in AI safety have increasingly emphasized the deep connection between AI safety and existential risk from advanced AI systems, suggesting that work on AI safety necessarily entails serious consideration of…

计算机与社会 · 计算机科学 2025-02-17 Balint Gyevnar , Atoosa Kasirzadeh

Artificial intelligence (AI) represents a technological upheaval with the potential to change human society. Because of its transformative potential, AI is increasingly becoming subject to regulatory initiatives at the global level. Yet, so…

综合经济学 · 经济学 2023-05-22 Jonas Tallberg , Eva Erman , Markus Furendal , Johannes Geith , Mark Klamberg , Magnus Lundgren

As frontier AI systems advance toward transformative capabilities, we need a parallel transformation in how we measure and evaluate these systems to ensure safety and inform governance. While benchmarks have been the primary method for…

人工智能 · 计算机科学 2025-05-12 Markov Grey , Charbel-Raphaël Segerie

Much of the research focus on AI alignment seeks to align large language models and other foundation models to the context-less and generic values of helpfulness, harmlessness, and honesty. Frontier model providers also strive to align…

计算机与社会 · 计算机科学 2025-01-23 Kush R. Varshney , Zahra Ashktorab , Djallel Bouneffouf , Matthew Riemer , Justin D. Weisz

Enabled and driven by modern advances in wireless telecommunication and artificial intelligence, the convergence of communication, computing, and control is becoming inevitable in future industrial applications. Analytical and optimizing…

系统与控制 · 电气工程与系统科学 2022-11-07 Bin Han , Hans D. Schotten

As AI systems become more capable and widely deployed as agents, ensuring their safe operation becomes critical. AI control offers one approach to mitigating the risk from untrusted AI agents by monitoring their actions and intervening or…

人工智能 · 计算机科学 2025-11-06 Jon Kutasov , Chloe Loughridge , Yuqi Sun , Henry Sleight , Buck Shlegeris , Tyler Tracy , Joe Benton

Value alignment is essential for building AI systems that can safely and reliably interact with people. However, what a person values -- and is even capable of valuing -- depends on the concepts that they are currently using to understand…

人工智能 · 计算机科学 2023-11-01 Sunayana Rane , Mark Ho , Ilia Sucholutsky , Thomas L. Griffiths

This comprehensive review paper provides a thorough examination of current advancements and research in the field of arc fault detection for electrical distribution systems. The increasing demand for electricity, coupled with the increasing…

AI alignment is often framed as the task of ensuring that an AI system follows a set of stated principles or human preferences, but general principles rarely determine their own application in concrete cases. When principles conflict, when…

人工智能 · 计算机科学 2026-04-14 Behrooz Razeghi

A morally acceptable course of AI development should avoid two dangers: creating unaligned AI systems that pose a threat to humanity and mistreating AI systems that merit moral consideration in their own right. This paper argues these two…

计算机与社会 · 计算机科学 2025-10-16 Adam Bradley , Bradford Saad
‹ 上一页 1 8 9 10 下一页 ›