English
Related papers

Related papers: Rationality is Self-Defeating in Permissionless Sy…

200 papers

Performing some task among a set of agents requires the use of some protocol that regulates the interactions between them. If those agents are rational, they may try to subvert the protocol for their own benefit, in an attempt to reach an…

Computer Science and Game Theory · Computer Science 2016-11-18 Josep Domingo-Ferrer , Jordi Soria-Comas , Oana Ciobotaru

Conversational implicatures are usually described as being licensed by the disobeying or flouting of a Principle of Cooperation. However, the specification of this principle has proved computationally elusive. In this paper we suggest that…

cmp-lg · Computer Science 2007-05-23 Mark Lee

Machines are being increasingly used in decision-making processes, resulting in the realization that decisions need explanations. Unfortunately, an increasing number of these deployed models are of a 'black-box' nature where the reasoning…

Artificial Intelligence · Computer Science 2023-11-07 Sopam Dasgupta

The off-switch problem is a critical challenge in AI control: if an AI system resists being switched off, it poses a significant risk. In this paper, we model the off-switch problem as a signalling game, where a human decision-maker…

Machine Learning · Computer Science 2025-04-01 Alessio Benavoli , Alessandro Facchini , Marco Zaffalon

Denial Logic DL, a system of justification logic, is the logic of an agent whose justified beliefs are false, who cannot avow his own propositional attitudes or believe tautologies, but who can believe contradictions. Using Artemov's…

Logic · Mathematics 2012-09-17 Florian Lengyel , Benoit St-Pierre

For an artificial intelligence (AI) to be aligned with human values (or human preferences), it must first learn those values. AI systems that are trained on human behavior, risk miscategorising human irrationalities as human values -- and…

Artificial Intelligence · Computer Science 2022-03-02 Rebecca Gorman , Stuart Armstrong

As artificial intelligence systems become increasingly agentic, capable of general reasoning, planning, and value prioritization, current safety practices that treat obedience as a proxy for ethical behavior are becoming inadequate. This…

Artificial Intelligence · Computer Science 2025-07-04 Joseph Boland

We present a comprehensive language theoretic causality analysis framework for explaining safety property violations in the setting of concurrent reactive systems. Our framework allows us to uniformly express a number of causality notions…

Formal Languages and Automata Theory · Computer Science 2019-01-04 Rayna Dimitrova , Rupak Majumdar , Vinayak S. Prabhu

Large Language Models (LLMs) are effective at deceiving, when prompted to do so. But under what conditions do they deceive spontaneously? Models that demonstrate better performance on reasoning tasks are also better at prompted deception.…

Computation and Language · Computer Science 2025-04-02 Samuel M. Taylor , Benjamin K. Bergen

This study investigates the relationship between resilience of control systems to attacks and the information available to malicious attackers. Specifically, it is shown that control systems are guaranteed to be secure in an asymptotic…

Systems and Control · Electrical Eng. & Systems 2020-03-27 Hampei Sasahara , Serkan Saritas , Henrik Sandberg

National and international guidelines for trustworthy artificial intelligence (AI) consider explainability to be a central facet of trustworthy systems. This paper outlines a multi-disciplinary rationale for explainability auditing.…

Computers and Society · Computer Science 2025-04-22 Markus Langer , Kevin Baum , Kathrin Hartmann , Stefan Hessel , Timo Speith , Jonas Wahl

Coding theory plays a crucial role in enabling reliable communication, storage, and computation. Classical approaches assume a worst-case adversarial model and ensure error correction and data recovery only when the number of honest nodes…

Machine Learning · Computer Science 2026-01-06 Hanzaleh Akbari Nodehi , Viveck R. Cadambe , Mohammad Ali Maddah-Ali

Assuming humans are (approximately) rational enables robots to infer reward functions by observing human behavior. But people exhibit a wide array of irrationalities, and our goal with this work is to better understand the effect they can…

Machine Learning · Computer Science 2021-11-16 Lawrence Chan , Andrew Critch , Anca Dragan

Latest insights from biology show that intelligence not only emerges from the connections between neurons but that individual neurons shoulder more computational responsibility than previously anticipated. This perspective should be…

Machine Learning · Computer Science 2024-03-19 Quentin Delfosse , Patrick Schramowski , Martin Mundt , Alejandro Molina , Kristian Kersting

The quality of rationales is essential in the reasoning capabilities of language models. Rationales not only enhance reasoning performance in complex natural language tasks but also justify model decisions. However, obtaining impeccable…

Computation and Language · Computer Science 2025-03-05 Hazel H. Kim

We discover a novel and surprising phenomenon of unintentional misalignment in reasoning language models (RLMs), which we call self-jailbreaking. Specifically, after benign reasoning training on math or code domains, RLMs will use multiple…

Cryptography and Security · Computer Science 2026-04-30 Zheng-Xin Yong , Stephen H. Bach

The main aim of decision support systems is to find solutions that satisfy user requirements. Often, this leads to predictability of those solutions, in the sense that having the input data and the model, an adversary or enemy can predict…

Discrete Mathematics · Computer Science 2021-01-18 Daniel Karapetyan , Andrew J. Parkes

With the availability of large datasets and ever-increasing computing power, there has been a growing use of data-driven artificial intelligence systems, which have shown their potential for successful application in diverse areas. However,…

Cryptography and Security · Computer Science 2021-08-05 Jose N. Paredes , Juan Carlos L. Teze , Gerardo I. Simari , Maria Vanina Martinez

An Artificially Intelligent system (an AI) has debatable personhood if it's epistemically possible either that the AI is a person or that it falls far short of personhood. Debatable personhood is a likely outcome of AI development and might…

Computers and Society · Computer Science 2023-03-31 Eric Schwitzgebel

Authentication systems are designed to give the right person access to an organization's information system and to restrict it from the wrong person. Such systems are designed by IT professionals to protect an organization's assets (e.g.,…

Cryptography and Security · Computer Science 2015-09-04 Christopher S. Pilson , James C. McElroy