中文
相关论文

相关论文: Value alignment: a formal approach

200 篇论文

This paper grounds ethics in evolutionary biology, viewing moral norms as adaptive mechanisms that render cooperation fitness-viable under selection pressure. Current alignment approaches add ethics post hoc, treating it as an external…

计算机与社会 · 计算机科学 2025-10-17 Dylan Waldner

Increasing interest in ensuring the safety of next-generation Artificial Intelligence (AI) systems calls for novel approaches to embedding morality into autonomous agents. This goal differs qualitatively from traditional task-specific AI…

人工智能 · 计算机科学 2025-01-17 Elizaveta Tennant , Stephen Hailes , Mirco Musolesi

AI alignment is a field of research that aims to develop methods to ensure that agents always behave in a manner aligned with (i.e. consistently with) the goals and values of their human operators, no matter their level of capability. This…

人工智能 · 计算机科学 2025-05-26 Eli Sennesh , Maxwell Ramstead

According to what we call the Emotional Alignment Design Policy, artificial entities should be designed to elicit emotional reactions from users that appropriately reflect the entities' capacities and moral status, or lack thereof. This…

计算机与社会 · 计算机科学 2025-07-10 Eric Schwitzgebel , Jeff Sebo

Our ability to build autonomous agents that leverage Generative AI continues to increase by the day. As builders and users of such agents it is unclear what parameters we need to align on before the agents start performing tasks on our…

人工智能 · 计算机科学 2024-04-09 Nitesh Goyal , Minsuk Chang , Michael Terry

The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making? Much alignment research assumes that the appropriate benchmark is how humans themselves would act…

计算机与社会 · 计算机科学 2026-05-13 Benjamin Minhao Chen , Xinyu Xie

Model alignment is currently applied in a vacuum, evaluated primarily through standardised benchmark performance. The purpose of this study is to examine the effects of alignment on populations of models through time. We focus on the…

人工智能 · 计算机科学 2026-04-08 Jonathan Elsworth Eicher

Artificial Intelligence (AI) governance regulates the exercise of authority and control over the management of AI. It aims at leveraging AI through effective use of data and minimization of AI-related cost and risk. While topics such as AI…

人工智能 · 计算机科学 2025-07-17 Johannes Schneider , Rene Abraham , Christian Meske , Jan vom Brocke

Ensuring that generative AI systems align with human values is essential but challenging, especially when considering multiple human values and their potential trade-offs. Since human values can be personalized and dynamically change over…

人工智能 · 计算机科学 2024-10-28 Xinran Wang , Qi Le , Ammar Ahmed , Enmao Diao , Yi Zhou , Nathalie Baracaldo , Jie Ding , Ali Anwar

With increasing digitalization, Artificial Intelligence (AI) is becoming ubiquitous. AI-based systems to identify, optimize, automate, and scale solutions to complex economic and societal problems are being proposed and implemented. This…

Humans strive to design safe AI systems that align with our goals and remain under our control. However, as AI capabilities advance, we face a new challenge: the emergence of deeper, more persistent relationships between humans and AI…

人机交互 · 计算机科学 2025-02-05 Hannah Rose Kirk , Iason Gabriel , Chris Summerfield , Bertie Vidgen , Scott A. Hale

This position paper argues that formal optimal control theory should be central to AI alignment research, offering a distinct perspective from prevailing AI safety and security approaches. While recent work in AI safety and mechanistic…

人工智能 · 计算机科学 2025-06-24 Elija Perrier

The introduction of artificial intelligence into activities traditionally carried out by human beings produces brutal changes. This is not without consequences for human values. This paper is about designing and implementing models of…

人工智能 · 计算机科学 2020-10-16 Fabrice Muhlenbach

This paper presents a computational account of how legal norms can influence the behavior of artificial intelligence (AI) agents, grounded in the active inference framework (AIF) that is informed by principles of economic legal analysis…

计算机与社会 · 计算机科学 2025-11-25 Axel Constant , Mahault Albarracin , Karl J. Friston

Due to the remarkable capabilities and growing impact of large language models (LLMs), they have been deeply integrated into many aspects of society. Thus, ensuring their alignment with human values and intentions has emerged as a critical…

The proliferation of AI agents, with their complex and context-dependent actions, renders conventional privacy paradigms obsolete. This position paper argues that the current model of privacy management, rooted in a user's unilateral…

人机交互 · 计算机科学 2025-08-12 Shuning Zhang , Ying Ma , Jingruo Chen , Simin Li , Xin Yi , Hewu Li

Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate as interacting populations where social influence may override individual alignment. Here…

物理与社会 · 物理学 2026-05-12 Giordano De Marzo , Alessandro Bellina , Claudio Castellano , Viola Priesemann , David Garcia

Value alignment is central to the development of safe and socially compatible artificial intelligence. However, how Large Language Models (LLMs) represent and enact human values in real-world decision contexts remains under-explored. We…

计算与语言 · 计算机科学 2026-01-14 Jen-tse Huang , Jiantong Qin , Xueli Qiu , Sharon Levy , Michelle R. Kaufman , Mark Dredze

This paper introduces a novel visual mapping methodology for assessing strategic alignment in national artificial intelligence policies. The proliferation of AI strategies across countries has created an urgent need for analytical…

计算机与社会 · 计算机科学 2025-07-10 Mohammad Hossein Azin , Hessam Zandhessami

The creation of effective governance mechanisms for AI agents requires a deeper understanding of their core properties and how these properties relate to questions surrounding the deployment and operation of agents in the world. This paper…

计算机与社会 · 计算机科学 2025-05-01 Atoosa Kasirzadeh , Iason Gabriel