English
Related papers

Related papers: Avoiding Obfuscation with Prover-Estimator Debate

200 papers

In democracies, major policy decisions typically require some form of majority or consensus, so elites must secure mass support to govern. Historically, elites could shape support only through limited instruments like schooling and mass…

General Economics · Economics 2025-12-04 Nadav Kunievsky

Persuasion is a fundamental aspect of communication, influencing decision-making across diverse contexts, from everyday conversations to high-stakes scenarios such as politics, marketing, and law. The rise of conversational AI systems has…

Computation and Language · Computer Science 2026-03-24 Nimet Beyza Bozdag , Shuhaib Mehri , Xiaocheng Yang , Hyeonjeong Ha , Zirui Cheng , Esin Durmus , Jiaxuan You , Heng Ji , Gokhan Tur , Dilek Hakkani-Tür

Artificial intelligence systems are being increasingly deployed due to their potential to increase the efficiency, scale, consistency, fairness, and accuracy of decisions. However, as many of these systems are opaque in their operation,…

Artificial intelligence (AI) has demonstrated strong potential in clinical diagnostics, often achieving accuracy comparable to or exceeding that of human experts. A key challenge, however, is that AI reasoning frequently diverges from…

Artificial Intelligence · Computer Science 2026-05-25 Belona Sonna , Alban Grastien

To make AI systems broadly useful for challenging real-world tasks, we need them to learn complex human goals and preferences. One approach to specifying complex goals asks humans to judge during training which agent behaviors are safe and…

Machine Learning · Statistics 2018-10-23 Geoffrey Irving , Paul Christiano , Dario Amodei

This paper argues that a range of current AI systems have learned how to deceive humans. We define deception as the systematic inducement of false beliefs in the pursuit of some outcome other than the truth. We first survey empirical…

Computers and Society · Computer Science 2023-08-29 Peter S. Park , Simon Goldstein , Aidan O'Gara , Michael Chen , Dan Hendrycks

Artificial Intelligence (AI) is being increasingly deployed in practical applications. However, there is a major concern whether AI systems will be trusted by humans. In order to establish trust in AI systems, there is a need for users to…

Artificial Intelligence · Computer Science 2021-02-16 Quratul-ain Mahesar , Simon Parsons

As the use of artificial intelligence (AI) in high-stakes decision-making increases, the ability to contest such decisions is being recognised in AI ethics guidelines as an important safeguard for individuals. Yet, there is little guidance…

Human-Computer Interaction · Computer Science 2021-02-23 Henrietta Lyons , Eduardo Velloso , Tim Miller

Computational argumentation offers formal frameworks for transparent, verifiable reasoning but has traditionally been limited by its reliance on domain-specific information and extensive feature engineering. In contrast, LLMs excel at…

Artificial Intelligence · Computer Science 2026-03-18 Stylianos Loukas Vasileiou , Antonio Rago , Francesca Toni , William Yeoh

While multi-agent debate has been proposed as a promising strategy for improving AI reasoning ability, we find that debate can sometimes be harmful rather than helpful. Prior work has primarily focused on debates within homogeneous groups…

Computation and Language · Computer Science 2025-10-14 Andrea Wynn , Harsh Satija , Gillian Hadfield

In the face of rapidly advancing AI technology, individuals will increasingly rely on AI agents to navigate life's growing complexities, raising critical concerns about maintaining both human agency and autonomy. This paper addresses a…

Computers and Society · Computer Science 2025-04-29 Philipp Koralus

Many real world learning tasks involve complex or hard-to-specify objectives, and using an easier-to-specify proxy can lead to poor performance or misaligned behavior. One solution is to have humans provide a training signal by…

Machine Learning · Computer Science 2018-10-22 Paul Christiano , Buck Shlegeris , Dario Amodei

Artificial Intelligence (AI) has been used extensively in automatic decision making in a broad variety of scenarios, ranging from credit ratings for loans to recommendations of movies. Traditional design guidelines for AI models focus…

Artificial Intelligence · Computer Science 2018-09-27 Marisa Vasconcelos , Carlos Cardonha , Bernardo Gonçalves

An optimal delivery of arguments is key to persuasion in any debate, both for humans and for AI systems. This requires the use of clear and fluent claims relevant to the given debate. Prior work has studied the automatic assessment of…

Computation and Language · Computer Science 2023-09-08 Gabriella Skitalinskaya , Maximilian Spliethöver , Henning Wachsmuth

Machine learning systems increasingly make life-changing decisions about individuals, such as loan approvals, hiring, and cheating detection, raising a pressing question: how can individuals respond to negative decisions made by these…

Machine Learning · Statistics 2026-05-18 Timo Freiesleben , Kristof Meding , Gunnar König

Equipping agents with the capacity to justify made decisions using supporting evidence represents a cornerstone of accountable decision-making. Furthermore, ensuring that justifications are in line with human expectations and societal norms…

Machine Learning · Computer Science 2024-02-27 Aleksa Sukovic , Goran Radanovic

An Artificially Intelligent system (an AI) has debatable personhood if it's epistemically possible either that the AI is a person or that it falls far short of personhood. Debatable personhood is a likely outcome of AI development and might…

Computers and Society · Computer Science 2023-03-31 Eric Schwitzgebel

Artificial Intelligence (AI) systems are increasingly deployed in legal contexts, where their opacity raises significant challenges for fairness, accountability, and trust. The so-called ``black box problem'' undermines the legitimacy of…

Artificial Intelligence · Computer Science 2025-10-14 Andrada Iulia Prajescu , Roberto Confalonieri

With AI systems becoming more powerful and pervasive, there is increasing debate about keeping their actions aligned with the broader goals and needs of humanity. This multi-disciplinary and multi-stakeholder debate must resolve many…

Artificial Intelligence · Computer Science 2021-12-21 Koen Holtman

Automated verbal deception detection using methods from Artificial Intelligence (AI) has been shown to outperform humans in disentangling lies from truths. Research suggests that transparency and interpretability of computational methods…

Human-Computer Interaction · Computer Science 2026-04-10 Riccardo Loconte , Merylin Monaro , Pietro Pietrini , Bruno Verschuere , Bennett Kleinberg