English
Related papers

Related papers: Quantifying Misalignment Between Agents: Towards a…

200 papers

We introduce and formalize misalignment, a phenomenon of interactive environments perceived from an analyst's perspective where an agent holds beliefs about another agent's beliefs that do not correspond to the actual beliefs of the latter.…

Theoretical Economics · Economics 2025-08-01 Pierfrancesco Guarino , Gabriel Ziegler

Social alignment in AI systems aims to ensure that these models behave according to established societal values. However, unlike humans, who derive consensus on value judgments through social interaction, current language models (LMs) are…

Computation and Language · Computer Science 2023-10-31 Ruibo Liu , Ruixin Yang , Chenyan Jia , Ge Zhang , Denny Zhou , Andrew M. Dai , Diyi Yang , Soroush Vosoughi

The rapid advancement of artificial intelligence (AI) systems suggests that artificial general intelligence (AGI) systems may soon arrive. Many researchers are concerned that AIs and AGIs will harm humans via intentional misuse (AI-misuse)…

Artificial Intelligence · Computer Science 2023-05-31 Catalin Mitelut , Ben Smith , Peter Vamplew

This paper explores the potential of a multidisciplinary approach to testing and aligning artificial intelligence (AI), specifically focusing on large language models (LLMs). Due to the rapid development and wide application of LLMs,…

Computers and Society · Computer Science 2025-01-07 Ljubisa Bojic , Matteo Cinelli , Dubravko Culibrk , Boris Delibasic

The proliferation of AI agents, with their complex and context-dependent actions, renders conventional privacy paradigms obsolete. This position paper argues that the current model of privacy management, rooted in a user's unilateral…

Human-Computer Interaction · Computer Science 2025-08-12 Shuning Zhang , Ying Ma , Jingruo Chen , Simin Li , Xin Yi , Hewu Li

Beneficial societal outcomes cannot be guaranteed by aligning individual AI systems with the intentions of their operators or users. Even an AI system that is perfectly aligned to the intentions of its operating organization can lead to bad…

The deployment of decision-making AI agents presents a critical challenge in maintaining alignment with human values or guidelines while operating in complex, dynamic environments. Agents trained solely to achieve their objectives may adopt…

Artificial Intelligence · Computer Science 2025-12-09 Dena Mujtaba , Brian Hu , Anthony Hoogs , Arslan Basharat

A core challenge in the development of increasingly capable AI systems is to make them safe and reliable by ensuring their behaviour is consistent with human values. This challenge, known as the alignment problem, does not merely apply to…

Machine Learning · Computer Science 2023-11-07 Raphaël Millière

The burgeoning integration of artificial intelligence (AI) into human society brings forth significant implications for societal governance and safety. While considerable strides have been made in addressing AI alignment challenges,…

Artificial Intelligence · Computer Science 2025-06-17 Zhaowei Zhang , Fengshuo Bai , Mingzhi Wang , Haoyang Ye , Chengdong Ma , Yaodong Yang

Recent advances in AI -- including generative approaches -- have resulted in technology that can support humans in scientific discovery and forming decisions, but may also disrupt democracies and target individuals. The responsible use of…

Much of the research focus on AI alignment seeks to align large language models and other foundation models to the context-less and generic values of helpfulness, harmlessness, and honesty. Frontier model providers also strive to align…

Computers and Society · Computer Science 2025-01-23 Kush R. Varshney , Zahra Ashktorab , Djallel Bouneffouf , Matthew Riemer , Justin D. Weisz

Value alignment has emerged in recent years as a basic principle to produce beneficial and mindful Artificial Intelligence systems. It mainly states that autonomous entities should behave in a way that is aligned with our human values. In…

Multiagent Systems · Computer Science 2021-06-28 Nieves Montes , Carles Sierra

Aligning AI systems with organizational decision-making is typically framed as a single-target problem: make the model behave like the organization. We argue this framing obscures a deeper pluralistic challenge. We rely on a decision-policy…

Artificial Intelligence · Computer Science 2026-05-26 Niklas Weller , Emilio Barkett

For artificial intelligence to be beneficial to humans the behaviour of AI agents needs to be aligned with what humans want. In this paper we discuss some behavioural issues for language agents, arising from accidental misspecification by…

Artificial Intelligence · Computer Science 2021-03-30 Zachary Kenton , Tom Everitt , Laura Weidinger , Iason Gabriel , Vladimir Mikulik , Geoffrey Irving

AI design characteristics and human personality traits each impact the quality and outcomes of human-AI interactions. However, their relative and joint impacts are underexplored in imperfectly cooperative scenarios, where people and AI only…

Computation and Language · Computer Science 2026-04-20 Myke C. Cohen , Mingqian Zheng , Neel Bhandari , Hsien-Te Kao , Xuhui Zhou , Daniel Nguyen , Laura Cassani , Maarten Sap , Svitlana Volkova

This paper introduces a novel visual mapping methodology for assessing strategic alignment in national artificial intelligence policies. The proliferation of AI strategies across countries has created an urgent need for analytical…

Computers and Society · Computer Science 2025-07-10 Mohammad Hossein Azin , Hessam Zandhessami

Cooperative perception has attracted wide attention given its capability to leverage shared information across connected automated vehicles (CAVs) and smart infrastructures to address sensing occlusion and range limitation issues. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Zonglin Meng , Yun Zhang , Zhaoliang Zheng , Zhihao Zhao , Jiaqi Ma

A key concern with the concept of "alignment" is the implicit question of "alignment to what?". AI systems are increasingly used across the world, yet safety alignment is often focused on homogeneous monolingual settings. Additionally,…

Computation and Language · Computer Science 2024-07-09 Aakanksha , Arash Ahmadian , Beyza Ermis , Seraphina Goldfarb-Tarrant , Julia Kreutzer , Marzieh Fadaee , Sara Hooker

There is no 'ordinary' when it comes to AI. The human-AI experience is extraordinarily complex and specific to each person, yet dominant measures such as usability scales and engagement metrics flatten away nuance. We argue for AI…

Human-Computer Interaction · Computer Science 2026-03-11 Bhada Yun , Evgenia Taranova , Dana Feng , Renn Su , April Yi Wang

Deep neural networks excel in medical imaging but remain prone to biases, leading to fairness gaps across demographic groups. We provide the first systematic exploration of Human-AI alignment and fairness in this domain. Our results show…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Haozhe Luo , Ziyu Zhou , Zixin Shu , Aurélie Pahud de Mortanges , Robert Berke , Mauricio Reyes
‹ Prev 1 3 4 5 6 7 10 Next ›