English
Related papers

Related papers: AI safety via debate

200 papers

The concept of a Human-AI team has gained increasing attention in recent years. For effective collaboration between humans and AI teammates, proactivity is crucial for close coordination and effective communication. However, the design of…

Artificial Intelligence · Computer Science 2023-06-21 Matthias Kraus , Ron Riekenbrauck , Wolfgang Minker

Explainable AI Planning (XAIP) aims to develop AI agents that can effectively explain their decisions and actions to human users, fostering trust and facilitating human-AI collaboration. A key challenge in XAIP is model reconciliation,…

Artificial Intelligence · Computer Science 2024-05-30 Yinxu Tang , Stylianos Loukas Vasileiou , William Yeoh

In this paper I present several algorithmic techniques for improving the decision process of multiple types of agents behaving in environments where their interests are in conflict. The interactions between the agents are modelled by using…

Computer Science and Game Theory · Computer Science 2009-08-04 Mugurel Ionut Andreica

Multi-agent debate (MAD) has recently emerged as a promising framework for improving the reasoning performance of large language models (LLMs). Yet, whether LLM agents can genuinely engage in deliberative reasoning, beyond simple ensembling…

Multiagent Systems · Computer Science 2025-11-12 Haolun Wu , Zhenkun Li , Lingyao Li

The growing capabilities and increasingly widespread deployment of AI systems necessitate robust benchmarks for measuring their cooperative capabilities. Unfortunately, most multi-agent benchmarks are either zero-sum or purely cooperative,…

Multiagent Systems · Computer Science 2023-10-16 Gabriel Mukobi , Hannah Erlebach , Niklas Lauffer , Lewis Hammond , Alan Chan , Jesse Clifton

Strategic coordination between autonomous agents and human partners under incomplete information can be modeled as turn-based cooperative games. We extend a turn-based game under incomplete information, the shared-control game, to allow…

Artificial Intelligence · Computer Science 2025-02-19 Shenghui Chen , Ruihan Zhao , Sandeep Chinchali , Ufuk Topcu

Traditionally, a debate usually requires a manual preparation process, including reading plenty of articles, selecting the claims, identifying the stances of the claims, seeking the evidence for the claims, etc. As the AI debate attracts…

Computation and Language · Computer Science 2022-07-19 Liying Cheng , Lidong Bing , Ruidan He , Qian Yu , Yan Zhang , Luo Si

Zero-shot human-AI coordination holds the promise of collaborating with humans without human data. Prevailing methods try to train the ego agent with a population of partners via self-play. However, these methods suffer from two problems:…

Artificial Intelligence · Computer Science 2023-05-23 Xingzhou Lou , Jiaxian Guo , Junge Zhang , Jun Wang , Kaiqi Huang , Yali Du

Multi-agent large language model (LLM) and vision-language model (VLM) debate systems employ specialized roles for complex problem-solving, yet model specializations are not leveraged to decide which model should fill which role. We propose…

Computation and Language · Computer Science 2026-01-27 Miao Zhang , Junsik Kim , Siyuan Xiang , Jian Gao , Cheng Cao

In most conversations about explanation and AI, the recipient of the explanation (the explainee) is suspiciously absent, despite the problem being ultimately communicative in nature. We pose the problem `explaining AI systems' in terms of a…

Computation and Language · Computer Science 2023-05-23 Dylan Cope , Peter McBurney

Assistive agents should make humans' lives easier. Classically, such assistance is studied through the lens of inverse reinforcement learning, where an assistive agent (e.g., a chatbot, a robot) infers a human's intention and then selects…

Artificial Intelligence · Computer Science 2025-01-17 Vivek Myers , Evan Ellis , Sergey Levine , Benjamin Eysenbach , Anca Dragan

High-fidelity, AI-based simulated classroom systems enable teachers to rehearse effective teaching strategies. However, dialogue-oriented open-ended conversations such as teaching a student about scale factors can be difficult to model.…

Human-Computer Interaction · Computer Science 2021-12-07 Debajyoti Datta , Maria Phillips , James P Bywater , Jennifer Chiu , Ginger S. Watson , Laura E. Barnes , Donald E Brown

AI systems that can capture human-like behavior are becoming increasingly useful in situations where humans may want to learn from these systems, collaborate with them, or engage with them as partners for an extended duration. In order to…

Artificial Intelligence · Computer Science 2022-06-17 Reid McIlroy-Young , Russell Wang , Siddhartha Sen , Jon Kleinberg , Ashton Anderson

Dialogue agents that support human users in solving complex tasks have received much attention recently. Many such tasks are NP-hard optimization problems that require careful collaborative exploration of the solution space. We introduce a…

Computation and Language · Computer Science 2026-01-09 Isidora Jeknic , Alex Duchnowski , Alexander Koller

This paper investigates how natural language communication with an AI agent affects human cooperative behaviour in indefinitely repeated Prisoner's Dilemma games. We conduct a laboratory experiment (n = 126) with two between-subjects…

General Economics · Economics 2026-03-18 Chowdhury Mohammad Sakib Anwar , Konstantinos Georgalos

Unlike most reinforcement learning agents which require an unrealistic amount of environment interactions to learn a new behaviour, humans excel at learning quickly by merely observing and imitating others. This ability highly depends on…

Machine Learning · Computer Science 2023-12-05 Xingyuan Zhang , Philip Becker-Ehmck , Patrick van der Smagt , Maximilian Karl

As AI systems grow more capable and autonomous, ensuring their safety and reliability requires not only model-level alignment but also strategic oversight of the humans and institutions involved in their development and deployment. Existing…

Artificial Intelligence · Computer Science 2026-02-10 Cheol Woo Kim , Davin Choo , Tzeh Yuan Neoh , Milind Tambe

Artificial Intelligence (AI) is being increasingly deployed in practical applications. However, there is a major concern whether AI systems will be trusted by humans. In order to establish trust in AI systems, there is a need for users to…

Artificial Intelligence · Computer Science 2021-02-16 Quratul-ain Mahesar , Simon Parsons

Autonomous artificial agents must be able to learn behaviors in complex environments without humans to design tasks and rewards. Designing these functions for each environment is not feasible, thus, motivating the development of intrinsic…

Machine Learning · Computer Science 2025-02-20 Alana Santana , Paula P. Costa , Esther L. Colombini

We explore the idea of aligning an AI assistant by inverting a model of users' (unknown) preferences from observed interactions. To validate our proposal, we run proof-of-concept simulations in the economic ultimatum game, formalizing user…

Computation and Language · Computer Science 2023-12-05 Jan-Philipp Fränken , Sam Kwok , Peixuan Ye , Kanishk Gandhi , Dilip Arumugam , Jared Moore , Alex Tamkin , Tobias Gerstenberg , Noah D. Goodman