中文
相关论文

相关论文: Aligning Artificial Superintelligence via a Multi-…

200 篇论文

We introduce AuditBench, an alignment auditing benchmark. AuditBench consists of 56 language models with implanted hidden behaviors. Each model has one of 14 concerning behaviors--such as sycophantic deference, opposition to AI regulation,…

计算与语言 · 计算机科学 2026-03-11 Abhay Sheshadri , Aidan Ewart , Kai Fronsdal , Isha Gupta , Samuel R. Bowman , Sara Price , Samuel Marks , Rowan Wang

Recent advances in general-purpose AI underscore the urgent need to align AI systems with human goals and values. Yet, the lack of a clear, shared understanding of what constitutes "alignment" limits meaningful progress and…

The value-alignment problem for artificial intelligence (AI) asks how we can ensure that the 'values' (i.e., objective functions) of artificial systems are aligned with the values of humanity. In this paper, I argue that linguistic…

人工智能 · 计算机科学 2022-07-05 Travis LaCroix

AI alignment research aims to develop techniques to ensure that AI systems do not cause harm. However, every alignment technique has failure modes, which are conditions in which there is a non-negligible chance that the technique fails to…

人工智能 · 计算机科学 2025-10-14 Leonard Dung , Florian Mai

Artificial agents now generate behavior rich enough to invite trust, surprise, and concern, yet our evaluation tools still privilege capability scores over psychological structure. This paper argues that the philosophical impasse between…

人工智能 · 计算机科学 2026-05-26 Alex Bogdan , Adrian de Valois-Franklin

This paper presents the Artificial Agency Program (AAP), a position and research agenda for building AI systems as reality embedded, resource-bounded agents whose development is driven by curiosity-as-learning-progress under physical and…

人工智能 · 计算机科学 2026-03-02 Richard Csaky

The potential risk of AI systems unintentionally embedding and reproducing bias has attracted the attention of machine learning practitioners and society at large. As policy makers are willing to set the standards of algorithms and AI…

人工智能 · 计算机科学 2020-03-17 Boris Ruf , Chaouki Boutharouite , Marcin Detyniecki

Social Explainable AI (SAI) is a new direction in artificial intelligence that emphasises decentralisation, transparency, social context, and focus on the human users. SAI research is still at an early stage. Consequently, it concentrates…

多智能体系统 · 计算机科学 2023-10-20 Damian Kurpiewski , Wojciech Jamroga , Teofil Sidoruk

As conversational AI systems become more realistic and widely deployed, users are increasingly uncertain about whether they are interacting with a human or an AI system. When AI identity is unclear, users may unwittingly share sensitive…

人机交互 · 计算机科学 2026-03-19 Anna Gausen , Sarenne Wallbridge , Hannah Rose Kirk , Jennifer Williams , Christopher Summerfield

As artificial intelligence (AI) systems are increasingly deployed, principles for ethical AI are also proliferating. Certification offers a method to both incentivize adoption of these principles and substantiate that they have been…

计算机与社会 · 计算机科学 2021-05-24 Peter Cihon , Moritz J. Kleinaltenkamp , Jonas Schuett , Seth D. Baum

Decisions made by various Artificial Intelligence (AI) systems greatly influence our day-to-day lives. With the increasing use of AI systems, it becomes crucial to know that they are fair, identify the underlying biases in their…

计算机与社会 · 计算机科学 2022-03-15 Avinash Agarwal , Harsh Agarwal , Nihaarika Agarwal

As AI agents increasingly operate in multi-agent environments, understanding their collective behavior becomes critical for predicting the dynamics of artificial societies. This study examines conformity, the tendency to align with group…

人工智能 · 计算机科学 2026-01-12 Alessandro Bellina , Giordano De Marzo , David Garcia

Large language models are increasingly deployed not as single assistants but as committees whose members deliberate and then vote or synthesize a decision. Such systems are often expected to be more robust than individual models. We show…

人工智能 · 计算机科学 2026-04-07 Hajime Shimao , Warut Khern-am-nuai , Sung Joo Kim

Across various applications, humans increasingly use black-box artificial intelligence (AI) systems without insight into these systems' reasoning. To counter this opacity, explainable AI (XAI) methods promise enhanced transparency and…

人机交互 · 计算机科学 2025-01-09 Philipp Spitzer , Joshua Holstein , Katelyn Morrison , Kenneth Holstein , Gerhard Satzger , Niklas Kühl

Generic AI auto-complete for message composition often fails to capture the nuance of personal identity, requiring editing. While harmless in low-stakes settings, for users of Augmentative and Alternative Communication (AAC) devices, who…

人机交互 · 计算机科学 2026-02-23 Tobias M. Weinberg , Ricardo E. Gonzalez Penuela , Stephanie Valencia , Thijs Roumen

Increasingly, laws are being proposed and passed by governments around the world to regulate Artificial Intelligence (AI) systems implemented into the public and private sectors. Many of these regulations address the transparency of AI…

计算机与社会 · 计算机科学 2022-07-05 Andrew Bell , Oded Nov , Julia Stoyanovich

The promise of human-AI teaming lies in humans and AI working together to achieve performance levels neither could accomplish alone. Effective communication between AI and humans is crucial for teamwork, enabling users to efficiently…

人机交互 · 计算机科学 2025-08-13 Tina Behzad , Nikolos Gurney , Ning Wang , David V. Pynadath

This paper addresses a synchronization problem that arises when a team of aerial robots (ARs) need to communicate while performing assigned tasks in a cooperative scenario. Each robot has a limited communication range and flies within a…

机器人学 · 计算机科学 2019-02-15 J. M. Díaz-Báñez , L. E. Caraballo , M. A. Lopez , S. Bereg , I. Maza , A. Ollero

We present a silent, self-stabilizing ranking protocol for the population protocol model of distributed computing, where agents interact in randomly chosen pairs to solve a common task. We are given $n$ anonymous agents, and the goal is to…

分布式、并行与集群计算 · 计算机科学 2025-04-15 Petra Berenbrink , Robert Elsässer , Thorsten Götte , Lukas Hintze , Dominik Kaaser

This paper reviews and proposes concerns in adopting, fielding, and maintaining artificial intelligence (AI) systems. While the AI community has made rapid progress, there are challenges in certifying AI systems. Using procedures from…

人工智能 · 计算机科学 2021-11-04 Erik Blasch , Junchi Bin , Zheng Liu