English
Related papers

Related papers: (When) Is Truth-telling Favored in AI Debate?

200 papers

As AI systems are used to answer more difficult questions and potentially help create new knowledge, judging the truthfulness of their outputs becomes more difficult and more important. How can we supervise unreliable experts, which have…

Artificial Intelligence · Computer Science 2023-11-16 Julian Michael , Salsabila Mahdi , David Rein , Jackson Petty , Julien Dirani , Vishakh Padmakumar , Samuel R. Bowman

Training powerful AI systems to exhibit desired behaviors hinges on the ability to provide accurate human supervision on increasingly complex tasks. A promising approach to this problem is to amplify human judgement by leveraging the power…

Artificial Intelligence · Computer Science 2025-06-17 Jonah Brown-Cohen , Geoffrey Irving , Georgios Piliouras

As AI grows more powerful, it will increasingly shape how we understand the world. But with this influence comes the risk of amplifying misinformation and deepening social divides-especially on consequential topics where factual accuracy…

If AI systems match or exceed human capabilities on a wide range of tasks, it may become difficult for humans to efficiently judge their actions -- making it hard to use human feedback to steer them towards desirable traits. One proposed…

Artificial Intelligence · Computer Science 2025-05-26 Marie Davidsen Buhl , Jacob Pfau , Benjamin Hilton , Geoffrey Irving

The emergence of pre-trained AI systems with powerful capabilities across a diverse and ever-increasing set of complex domains has raised a critical challenge for AI safety as tasks can become too complicated for humans to judge directly.…

Artificial Intelligence · Computer Science 2023-11-27 Jonah Brown-Cohen , Geoffrey Irving , Georgios Piliouras

The core premise of AI debate as a scalable oversight technique is that it is harder to lie convincingly than to refute a lie, enabling the judge to identify the correct position. Yet, existing debate experiments have relied on datasets…

Scalable oversight protocols aim to enable humans to accurately supervise superhuman AI. In this paper we study debate, where two AI's compete to convince a judge; consultancy, where a single AI tries to convince a judge that asks…

Current QA systems can generate reasonable-sounding yet false answers without explanation or evidence for the generated answer, which is especially problematic when humans cannot readily check the model's answers. This presents a challenge…

Computation and Language · Computer Science 2022-04-14 Alicia Parrish , Harsh Trivedi , Ethan Perez , Angelica Chen , Nikita Nangia , Jason Phang , Samuel R. Bowman

In many contexts, lying -- the use of verbal falsehoods to deceive -- is harmful. While lying has traditionally been a human affair, AI systems that make sophisticated verbal statements are becoming increasingly prevalent. This raises the…

Computers and Society · Computer Science 2021-10-14 Owain Evans , Owen Cotton-Barratt , Lukas Finnveden , Adam Bales , Avital Balwit , Peter Wills , Luca Righetti , William Saunders

As Artificial Intelligence (AI) technology gets more intertwined with every system, people are using AI to make decisions on their everyday activities. In simple contexts, such as Netflix recommendations, or in more complex context like in…

Human-Computer Interaction · Computer Science 2020-03-04 Juliana Jansen Ferreira , Mateus de Souza Monteiro

As AI technologies are rolled out into healthcare, academia, human resources, law, and a multitude of other domains, they become de-facto arbiters of truth. But truth is highly contested, with many different definitions and approaches. This…

Computers and Society · Computer Science 2023-01-31 Luke Munn , Liam Magee , Vanicka Arora

While multi-agent debate has been proposed as a promising strategy for improving AI reasoning ability, we find that debate can sometimes be harmful rather than helpful. Prior work has primarily focused on debates within homogeneous groups…

Computation and Language · Computer Science 2025-10-14 Andrea Wynn , Harsh Satija , Gillian Hadfield

As the use of artificial intelligence (AI) in high-stakes decision-making increases, the ability to contest such decisions is being recognised in AI ethics guidelines as an important safeguard for individuals. Yet, there is little guidance…

Human-Computer Interaction · Computer Science 2021-02-23 Henrietta Lyons , Eduardo Velloso , Tim Miller

Common methods for aligning large language models (LLMs) with desired behaviour heavily rely on human-labelled data. However, as models grow increasingly sophisticated, they will surpass human expertise, and the role of human evaluation…

With AI systems becoming more powerful and pervasive, there is increasing debate about keeping their actions aligned with the broader goals and needs of humanity. This multi-disciplinary and multi-stakeholder debate must resolve many…

Artificial Intelligence · Computer Science 2021-12-21 Koen Holtman

Political online participation in the form of discussing political issues and exchanging opinions among citizens is gaining importance with more and more formats being held digitally. To come to a decision, a thorough discussion and…

Computation and Language · Computer Science 2026-03-27 Maike Behrendt , Stefan Sylvius Wagner , Carina Weinmann , Marike Bormann , Mira Warne , Stefan Harmeling

Common methods for aligning already-capable models with desired behavior rely on the ability of humans to provide supervision. However, future superhuman models will surpass the capability of humans. Therefore, humans will only be able to…

Computation and Language · Computer Science 2025-01-24 Hao Lang , Fei Huang , Yongbin Li

To make AI systems broadly useful for challenging real-world tasks, we need them to learn complex human goals and preferences. One approach to specifying complex goals asks humans to judge during training which agent behaviors are safe and…

Machine Learning · Statistics 2018-10-23 Geoffrey Irving , Paul Christiano , Dario Amodei

An Artificially Intelligent system (an AI) has debatable personhood if it's epistemically possible either that the AI is a person or that it falls far short of personhood. Debatable personhood is a likely outcome of AI development and might…

Computers and Society · Computer Science 2023-03-31 Eric Schwitzgebel

Recently there are increasing concerns about the fairness of Artificial Intelligence (AI) in real-world applications such as computer vision and recommendations. For example, recognition algorithms in computer vision are unfair to black…

Computation and Language · Computer Science 2020-11-03 Haochen Liu , Jamell Dacon , Wenqi Fan , Hui Liu , Zitao Liu , Jiliang Tang
‹ Prev 1 2 3 10 Next ›