English
Related papers

Related papers: Avoiding Obfuscation with Prover-Estimator Debate

200 papers

Given that Artificial Intelligence (AI) increasingly permeates our lives, it is critical that we systematically align AI objectives with the goals and values of humans. The human-AI alignment problem stems from the impracticality of…

Computers and Society · Computer Science 2022-07-05 John Nay , James Daily

AI agents are able to tackle increasingly complex tasks. To achieve more ambitious goals, AI agents need to be able to meaningfully decompose problems into manageable sub-components, and safely delegate their completion across to other AI…

Artificial Intelligence · Computer Science 2026-02-13 Nenad Tomašev , Matija Franklin , Simon Osindero

Artificial intelligence functions not as an epistemic leveller, but as an accelerant of cognitive stratification, entrenching and formalising informational castes within liberal-democratic societies. Synthesising formal epistemology,…

Computers and Society · Computer Science 2025-07-22 Craig S Wright

We conducted an International AI Negotiation Competition in which participants designed and refined prompts for AI negotiation agents. We then facilitated over 180,000 negotiations between these agents across multiple scenarios with diverse…

Artificial Intelligence · Computer Science 2026-01-15 Michelle Vaccaro , Michael Caosun , Harang Ju , Sinan Aral , Jared R. Curhan

Without the ability to estimate and benchmark AI capability advancements, organizations are left to respond to each change reactively, impeding their ability to build viable mid and long-term strategies. This paper explores the recent…

Computers and Society · Computer Science 2023-04-03 Emily Dardaman , Abhishek Gupta

Concerns about the risks and harms posed by artificial intelligence (AI) have resulted in significant study into algorithmic transparency, giving rise to a sub-field known as Explainable AI (XAI). Unfortunately, despite a decade of…

Computers and Society · Computer Science 2024-12-23 Andrew Bell , Julia Stoyanovich

Understanding when and why to apply any given eXplainable Artificial Intelligence (XAI) technique is not a straightforward task. There is no single approach that is best suited for a given context. This paper aims to address the challenge…

Artificial Intelligence · Computer Science 2023-12-14 Leila Methnani , Virginia Dignum , Andreas Theodorou

Explainability and its emerging counterpart contestability have become important normative and design principles for trustworthy AI as they enable users and subjects to understand and challenge AI decisions. However, realizing these…

Computers and Society · Computer Science 2025-08-15 Timothée Schmude , Mireia Yurrita , Kars Alfrink , Thomas Le Goff , Sebastian Tschiatschek , Tiphaine Viard

Machine learning systems perform well on pattern matching tasks, but their ability to perform algorithmic or logical reasoning is not well understood. One important reasoning capability is algorithmic extrapolation, in which models trained…

Machine Learning · Computer Science 2022-10-18 Arpit Bansal , Avi Schwarzschild , Eitan Borgnia , Zeyad Emam , Furong Huang , Micah Goldblum , Tom Goldstein

Current bioacoustic AI systems achieve impressive cross-species performance by processing animal communication through transformer architectures, foundation model paradigms, and other computational approaches. However, these approaches…

Artificial Intelligence · Computer Science 2025-11-13 Graham L. Bishop

Whether and how data scientists, statisticians and modellers should be accountable for the AI systems they develop remains a controversial and highly debated topic, especially given the complexity of AI systems and the difficulties in…

Artificial Intelligence · Computer Science 2023-09-12 Cassandra Bird , Daniel Williamson , Sabina Leonelli

This paper presents Abduction and Argumentation as two principled forms for reasoning, and fleshes out the fundamental role that they can play within Machine Learning. It reviews the state-of-the-art work over the past few decades on the…

Artificial Intelligence · Computer Science 2020-10-27 Antonis Kakas , Loizos Michael

Explainable Artificial Intelligence and Formal Argumentation have received significant attention in recent years. Argumentation-based systems often lack explainability while supporting decision-making processes. Counterfactual and…

Artificial Intelligence · Computer Science 2024-05-08 Gianvincenzo Alfano , Sergio Greco , Francesco Parisi , Irina Trubitsyna

In open-domain question answering, due to the ambiguity of questions, multiple plausible answers may exist. To provide feasible answers to an ambiguous question, one approach is to directly predict all valid answers, but this can struggle…

Computation and Language · Computer Science 2023-07-11 Weiwei Sun , Hengyi Cai , Hongshen Chen , Pengjie Ren , Zhumin Chen , Maarten de Rijke , Zhaochun Ren

Recent advances in AI reasoning models provide unprecedented transparency into their decision-making processes, transforming them from traditional black-box systems into models that articulate step-by-step chains of thought rather than…

Software Engineering · Computer Science 2025-03-04 Christoph Treude , Raula Gaikovina Kula

As artificial intelligence (AI) improves, traditional alignment strategies may falter in the face of unpredictable self-improvement, hidden subgoals, and the sheer complexity of intelligent systems. Inspired by contemplative wisdom…

Artificial Intelligence · Computer Science 2025-08-19 Ruben Laukkonen , Fionn Inglis , Shamil Chandaria , Lars Sandved-Smith , Edmundo Lopez-Sola , Jakob Hohwy , Jonathan Gold , Adam Elwood

The performance of adversarial dialogue generation models relies on the quality of the reward signal produced by the discriminator. The reward signal from a poor discriminator can be very sparse and unstable, which may lead the generator to…

Computation and Language · Computer Science 2018-12-11 Ziming Li , Julia Kiseleva , Maarten de Rijke

The use of artificial intelligence (AI) in the public sector is best understood as a continuation and intensification of long standing rationalization and bureaucratization processes. Drawing on Weber, we take the core of these processes to…

Artificial Intelligence · Computer Science 2024-07-09 Jakob Mokander , Ralph Schroeder

In contrast to software reverse engineering, there are hardly any tools available that support hardware reversing. Therefore, the reversing process is conducted by human analysts combining several complex semi-automated steps. However,…

Cryptography and Security · Computer Science 2019-10-02 Carina Wiesen , Nils Albartus , Max Hoffmann , Steffen Becker , Sebastian Wallat , Marc Fyrbiak , Nikol Rummel , Christof Paar

Generative AI systems increasingly expose powerful reasoning and image refinement capabilities through user-facing chatbot interfaces. In this work, we show that the na\"ive exposure of such capabilities fundamentally undermines modern…

Cryptography and Security · Computer Science 2026-03-12 Sunpill Kim , Chanwoo Hwang , Minsu Kim , Jae Hong Seo