English
Related papers

Related papers: AI safety via debate

200 papers

As AI usage becomes more prevalent in social contexts, understanding agent-user interaction is critical to designing systems that improve both individual and group outcomes. We present an online behavioral experiment (N = 243) in which…

Computer Science and Game Theory · Computer Science 2026-02-16 Kehang Zhu , Nithum Thain , Vivian Tsai , James Wexler , Crystal Qian

The emergence of Large Language Models (LLMs), has opened exciting possibilities for constructing computational simulations designed to replicate human behavior accurately. Current research suggests that LLM-based agents become increasingly…

Computation and Language · Computer Science 2024-12-18 Amir Taubenfeld , Yaniv Dover , Roi Reichart , Ariel Goldstein

As the use of artificial intelligence (AI) in high-stakes decision-making increases, the ability to contest such decisions is being recognised in AI ethics guidelines as an important safeguard for individuals. Yet, there is little guidance…

Human-Computer Interaction · Computer Science 2021-02-23 Henrietta Lyons , Eduardo Velloso , Tim Miller

In the same way that generative models today conduct most of their training in a self-supervised fashion, how can agentic models conduct their training in a self-supervised fashion, interactively exploring, learning, and preparing to…

Machine Learning · Computer Science 2025-10-21 Kathryn Wantlin , Chongyi Zheng , Benjamin Eysenbach

Achieving human-AI alignment in complex multi-agent games is crucial for creating trustworthy AI agents that enhance gameplay. We propose a method to evaluate this alignment using an interpretable task-sets framework, focusing on high-level…

Artificial Intelligence · Computer Science 2024-06-21 Sugandha Sharma , Guy Davidson , Khimya Khetarpal , Anssi Kanervisto , Udit Arora , Katja Hofmann , Ida Momennejad

As complex societal issues continue to emerge, fostering democratic skills like valuing diverse perspectives and collaborative decision-making is increasingly vital in education. In this paper, we propose a Peer Agent (PA) system designed…

Human-Computer Interaction · Computer Science 2025-08-13 Kyuwon Kim , Jaeryeong Hwang , Younseo Lee , Jeanhee Lee , Sung-Eun Kim , Hyo-Jeong So

Optimization of parameterized policies for reinforcement learning (RL) is an important and challenging problem in artificial intelligence. Among the most common approaches are algorithms based on gradient ascent of a score function…

Machine Learning · Computer Science 2020-06-15 Sriram Srinivasan , Marc Lanctot , Vinicius Zambaldi , Julien Perolat , Karl Tuyls , Remi Munos , Michael Bowling

This paper proposes an intent-aware multi-agent planning framework as well as a learning algorithm. Under this framework, an agent plans in the goal space to maximize the expected utility. The planning process takes the belief of other…

Artificial Intelligence · Computer Science 2018-03-07 Siyuan Qi , Song-Chun Zhu

Artificial intelligence (AI) has enabled agents to master complex video games, from first-person shooters like Counter-Strike to real-time strategy games such as StarCraft II and racing games like Gran Turismo. While these achievements are…

The goal of building dialogue agents that can converse with humans naturally has been a long-standing dream of researchers since the early days of artificial intelligence. The well-known Turing Test proposed to judge the ultimate validity…

Artificial Intelligence · Computer Science 2022-12-13 Tom Young

Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate single agents, leaving multi-agent risks such as coordination failure and conflict…

Artificial Intelligence · Computer Science 2026-05-25 Pepijn Cobben , Xuanqiang Angelo Huang , Thao Amelia Pham , Isabel Dahlgren , Terry Jingchen Zhang , Zhijing Jin

Human-like agents are an increasingly important topic in games and beyond. Believable non-player characters enhance the gaming experience by improving immersion and providing entertainment. They also offer players the opportunity to engage…

Artificial Intelligence · Computer Science 2025-06-11 Maciej Swiechowski , Dominik Slezak

Many emerging applications of AI--from scientific discovery to medical diagnosis--require agents to seek information strategically: forming hypotheses, asking targeted questions, and making decisions under uncertainty. In high-stakes…

Computation and Language · Computer Science 2026-03-09 Gabriel Grand , Valerio Pepe , Jacob Andreas , Joshua B. Tenenbaum

A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue that, even when research agents are not scheming to…

Artificial Intelligence · Computer Science 2026-05-18 Aleksandr Bowkis , Marie Davidsen Buhl , Jacob Pfau , Geoffrey Irving

An optimal delivery of arguments is key to persuasion in any debate, both for humans and for AI systems. This requires the use of clear and fluent claims relevant to the given debate. Prior work has studied the automatic assessment of…

Computation and Language · Computer Science 2023-09-08 Gabriella Skitalinskaya , Maximilian Spliethöver , Henning Wachsmuth

Machine learning models are being increasingly deployed to take, or assist in taking, complicated and high-impact decisions, from quasi-autonomous vehicles to clinical decision support systems. This poses challenges, particularly when…

Machine Learning · Computer Science 2023-11-14 Alex J. Chan , Alihan Huyuk , Mihaela van der Schaar

We obtain global, non-asymptotic convergence guarantees for independent learning algorithms in competitive reinforcement learning settings with two agents (i.e., zero-sum stochastic games). We consider an episodic setting where in each…

Machine Learning · Computer Science 2021-01-13 Constantinos Daskalakis , Dylan J. Foster , Noah Golowich

The development and popularization of large language models (LLMs) have raised concerns that they will be used to create tailor-made, convincing arguments to push false or misleading narratives online. Early work has found that language…

Computers and Society · Computer Science 2025-05-21 Francesco Salvi , Manoel Horta Ribeiro , Riccardo Gallotti , Robert West

While state-of-the-art language models have achieved impressive results, they remain susceptible to inference-time adversarial attacks, such as adversarial prompts generated by red teams arXiv:2209.07858. One approach proposed to improve…

Computation and Language · Computer Science 2024-01-12 Steffi Chern , Zhen Fan , Andy Liu

In large systems, it is important for agents to learn to act effectively, but sophisticated multi-agent learning algorithms generally do not scale. An alternative approach is to find restricted classes of games where simple, efficient…

Multiagent Systems · Computer Science 2009-03-16 Ian A. Kash , Eric J. Friedman , Joseph Y. Halpern