English
Related papers

Related papers: Counterfactual Planning in AGI Systems

200 papers

Explainable AI (XAI) is a research area whose objective is to increase trustworthiness and to enlighten the hidden mechanism of opaque machine learning techniques. This becomes increasingly important in case such models are applied to the…

Machine Learning · Computer Science 2021-04-19 Danilo Numeroso , Davide Bacciu

A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue that, even when research agents are not scheming to…

Artificial Intelligence · Computer Science 2026-05-18 Aleksandr Bowkis , Marie Davidsen Buhl , Jacob Pfau , Geoffrey Irving

Many modern robotics applications require robots to function autonomously in dynamic environments including other decision making agents, such as people or other robots. This calls for fast and scalable interactive motion planning. This…

Robotics · Computer Science 2016-10-27 A. Bordallo , F. Previtali , N. Nardelli , S. Ramamoorthy

There has been a recent resurgence of interest in explainable artificial intelligence (XAI) that aims to reduce the opaqueness of AI-based decision-making systems, allowing humans to scrutinize and trust them. Prior work in this context has…

Artificial Intelligence · Computer Science 2021-06-24 Sainyam Galhotra , Romila Pradhan , Babak Salimi

Deep neural networks (DNNs) can accurately decode task-related information from brain activations. However, because of the nonlinearity of the DNN, the decisions made by DNNs are hardly interpretable. One of the promising approaches for…

Neurons and Cognition · Quantitative Biology 2021-10-29 Teppei Matsui , Masato Taki , Trung Quang Pham , Junichi Chikazoe , Koji Jimura

We propose an interactive methodology for generating counterfactual explanations for univariate time series data in classification tasks by leveraging 2D projections and decision boundary maps to tackle interpretability challenges. Our…

Machine Learning · Computer Science 2024-08-21 Udo Schlegel , Julius Rauscher , Daniel A. Keim

Post-hoc explanations of machine learning models are crucial for people to understand and act on algorithmic predictions. An intriguing class of explanations is through counterfactuals, hypothetical examples that show people how to obtain a…

Machine Learning · Computer Science 2019-12-09 Ramaravind Kommiya Mothilal , Amit Sharma , Chenhao Tan

Algorithms are commonly used to predict outcomes under a particular decision or intervention, such as predicting whether an offender will succeed on parole if placed under minimal supervision. Generally, to learn such counterfactual…

Machine Learning · Statistics 2021-04-19 Amanda Coston , Edward H. Kennedy , Alexandra Chouldechova

Explainable Artificial Intelligence (XAI) is a pivotal research domain aimed at understanding the operational mechanisms of AI systems, particularly those considered ``black boxes'' due to their complex, opaque nature. XAI seeks to make…

Machine Learning · Computer Science 2024-05-22 José Daniel Pascual-Triana , Alberto Fernández , Javier Del Ser , Francisco Herrera

In the pursuit of artificial general intelligence (AGI), we tackle Abstraction and Reasoning Corpus (ARC) tasks using a novel two-pronged approach. We employ the Decision Transformer in an imitation learning paradigm to model human…

Artificial Intelligence · Computer Science 2023-06-16 Jaehyun Park , Jaegyun Im , Sanha Hwang , Mintaek Lim , Sabina Ualibekova , Sejin Kim , Sundong Kim

To safely interact with humans, AI agents must both know our norms and consider them during planning. However, such norm-guided planning has been less explored, only within communities of artificial agents, and has ignored the dynamic…

Artificial Intelligence · Computer Science 2026-05-28 Taylor Olson , Roberto Salas-Damian , Kenneth D. Forbus

Machine learning based decision making systems applied in safety critical areas require reliable high certainty predictions. For this purpose, the system can be extended by an reject option which allows the system to reject inputs where…

Machine Learning · Computer Science 2022-07-06 André Artelt , Barbara Hammer

Counterfactual Thinking is a human cognitive ability studied in a wide variety of domains. It captures the process of reasoning about a past event that did not occur, namely what would have happened had this event occurred, or, otherwise,…

Artificial Intelligence · Computer Science 2019-12-20 Luis Moniz Pereira , Francisco C. Santos

This paper develops a control-theoretic framework for analyzing agentic systems embedded within feedback control loops, where an AI agent may adapt controller parameters, select among control strategies, invoke external tools, reconfigure…

Systems and Control · Electrical Eng. & Systems 2026-03-26 Ali Eslami , Jiangbo Yu

The increasing application of Artificial Intelligence and Machine Learning models poses potential risks of unfair behavior and, in light of recent regulations, has attracted the attention of the research community. Several researchers…

Machine Learning · Computer Science 2023-02-17 Giandomenico Cornacchia , Vito Walter Anelli , Fedelucio Narducci , Azzurra Ragone , Eugenio Di Sciascio

We consider counterfactual explanations, the problem of minimally adjusting features in a source input instance so that it is classified as a target class under a given classifier. This has become a topic of recent interest as a way to…

Machine Learning · Computer Science 2021-03-02 Miguel Á. Carreira-Perpiñán , Suryabhan Singh Hada

Explainable AI is an important area of research within which Explainable Planning is an emerging topic. In this paper, we argue that Explainable Planning can be designed as a service -- that is, as a wrapper around an existing planning…

Artificial Intelligence · Computer Science 2019-08-15 Michael Cashmore , Anna Collins , Benjamin Krarup , Senka Krivic , Daniele Magazzeni , David Smith

Cybersecurity is being fundamentally reshaped by foundation-model-based artificial intelligence. Large language models now enable autonomous planning, tool orchestration, and strategic adaptation at scale, challenging security architectures…

Cryptography and Security · Computer Science 2025-12-30 Tao Li , Quanyan Zhu

Online hate speech has become increasingly prevalent on social media, causing harm to individuals and society. While automated content moderation has received considerable attention, user-driven counterspeech remains a less explored yet…

Motivation: Many high-performance DTA models have been proposed, but they are mostly black-box and thus lack human interpretability. Explainable AI (XAI) can make DTA models more trustworthy, and can also enable scientists to distill…

Artificial Intelligence · Computer Science 2021-06-03 Tri Minh Nguyen , Thomas P Quinn , Thin Nguyen , Truyen Tran