English
Related papers

Related papers: Aligning Artificial Superintelligence via a Multi-…

200 papers

We introduce AuditBench, an alignment auditing benchmark. AuditBench consists of 56 language models with implanted hidden behaviors. Each model has one of 14 concerning behaviors--such as sycophantic deference, opposition to AI regulation,…

Computation and Language · Computer Science 2026-03-11 Abhay Sheshadri , Aidan Ewart , Kai Fronsdal , Isha Gupta , Samuel R. Bowman , Sara Price , Samuel Marks , Rowan Wang

Recent advances in general-purpose AI underscore the urgent need to align AI systems with human goals and values. Yet, the lack of a clear, shared understanding of what constitutes "alignment" limits meaningful progress and…

The value-alignment problem for artificial intelligence (AI) asks how we can ensure that the 'values' (i.e., objective functions) of artificial systems are aligned with the values of humanity. In this paper, I argue that linguistic…

Artificial Intelligence · Computer Science 2022-07-05 Travis LaCroix

AI alignment research aims to develop techniques to ensure that AI systems do not cause harm. However, every alignment technique has failure modes, which are conditions in which there is a non-negligible chance that the technique fails to…

Artificial Intelligence · Computer Science 2025-10-14 Leonard Dung , Florian Mai

Artificial agents now generate behavior rich enough to invite trust, surprise, and concern, yet our evaluation tools still privilege capability scores over psychological structure. This paper argues that the philosophical impasse between…

Artificial Intelligence · Computer Science 2026-05-26 Alex Bogdan , Adrian de Valois-Franklin

This paper presents the Artificial Agency Program (AAP), a position and research agenda for building AI systems as reality embedded, resource-bounded agents whose development is driven by curiosity-as-learning-progress under physical and…

Artificial Intelligence · Computer Science 2026-03-02 Richard Csaky

The potential risk of AI systems unintentionally embedding and reproducing bias has attracted the attention of machine learning practitioners and society at large. As policy makers are willing to set the standards of algorithms and AI…

Artificial Intelligence · Computer Science 2020-03-17 Boris Ruf , Chaouki Boutharouite , Marcin Detyniecki

Social Explainable AI (SAI) is a new direction in artificial intelligence that emphasises decentralisation, transparency, social context, and focus on the human users. SAI research is still at an early stage. Consequently, it concentrates…

Multiagent Systems · Computer Science 2023-10-20 Damian Kurpiewski , Wojciech Jamroga , Teofil Sidoruk

As conversational AI systems become more realistic and widely deployed, users are increasingly uncertain about whether they are interacting with a human or an AI system. When AI identity is unclear, users may unwittingly share sensitive…

Human-Computer Interaction · Computer Science 2026-03-19 Anna Gausen , Sarenne Wallbridge , Hannah Rose Kirk , Jennifer Williams , Christopher Summerfield

As artificial intelligence (AI) systems are increasingly deployed, principles for ethical AI are also proliferating. Certification offers a method to both incentivize adoption of these principles and substantiate that they have been…

Computers and Society · Computer Science 2021-05-24 Peter Cihon , Moritz J. Kleinaltenkamp , Jonas Schuett , Seth D. Baum

Decisions made by various Artificial Intelligence (AI) systems greatly influence our day-to-day lives. With the increasing use of AI systems, it becomes crucial to know that they are fair, identify the underlying biases in their…

Computers and Society · Computer Science 2022-03-15 Avinash Agarwal , Harsh Agarwal , Nihaarika Agarwal

As AI agents increasingly operate in multi-agent environments, understanding their collective behavior becomes critical for predicting the dynamics of artificial societies. This study examines conformity, the tendency to align with group…

Artificial Intelligence · Computer Science 2026-01-12 Alessandro Bellina , Giordano De Marzo , David Garcia

Large language models are increasingly deployed not as single assistants but as committees whose members deliberate and then vote or synthesize a decision. Such systems are often expected to be more robust than individual models. We show…

Artificial Intelligence · Computer Science 2026-04-07 Hajime Shimao , Warut Khern-am-nuai , Sung Joo Kim

Across various applications, humans increasingly use black-box artificial intelligence (AI) systems without insight into these systems' reasoning. To counter this opacity, explainable AI (XAI) methods promise enhanced transparency and…

Human-Computer Interaction · Computer Science 2025-01-09 Philipp Spitzer , Joshua Holstein , Katelyn Morrison , Kenneth Holstein , Gerhard Satzger , Niklas Kühl

Generic AI auto-complete for message composition often fails to capture the nuance of personal identity, requiring editing. While harmless in low-stakes settings, for users of Augmentative and Alternative Communication (AAC) devices, who…

Human-Computer Interaction · Computer Science 2026-02-23 Tobias M. Weinberg , Ricardo E. Gonzalez Penuela , Stephanie Valencia , Thijs Roumen

Increasingly, laws are being proposed and passed by governments around the world to regulate Artificial Intelligence (AI) systems implemented into the public and private sectors. Many of these regulations address the transparency of AI…

Computers and Society · Computer Science 2022-07-05 Andrew Bell , Oded Nov , Julia Stoyanovich

The promise of human-AI teaming lies in humans and AI working together to achieve performance levels neither could accomplish alone. Effective communication between AI and humans is crucial for teamwork, enabling users to efficiently…

Human-Computer Interaction · Computer Science 2025-08-13 Tina Behzad , Nikolos Gurney , Ning Wang , David V. Pynadath

This paper addresses a synchronization problem that arises when a team of aerial robots (ARs) need to communicate while performing assigned tasks in a cooperative scenario. Each robot has a limited communication range and flies within a…

Robotics · Computer Science 2019-02-15 J. M. Díaz-Báñez , L. E. Caraballo , M. A. Lopez , S. Bereg , I. Maza , A. Ollero

We present a silent, self-stabilizing ranking protocol for the population protocol model of distributed computing, where agents interact in randomly chosen pairs to solve a common task. We are given $n$ anonymous agents, and the goal is to…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-04-15 Petra Berenbrink , Robert Elsässer , Thorsten Götte , Lukas Hintze , Dominik Kaaser

This paper reviews and proposes concerns in adopting, fielding, and maintaining artificial intelligence (AI) systems. While the AI community has made rapid progress, there are challenges in certifying AI systems. Using procedures from…

Artificial Intelligence · Computer Science 2021-11-04 Erik Blasch , Junchi Bin , Zheng Liu