English

Causal Analysis of Agent Behavior for AI Safety

Artificial Intelligence 2021-03-09 v1 Machine Learning

Abstract

As machine learning systems become more powerful they also become increasingly unpredictable and opaque. Yet, finding human-understandable explanations of how they work is essential for their safe deployment. This technical report illustrates a methodology for investigating the causal mechanisms that drive the behaviour of artificial agents. Six use cases are covered, each addressing a typical question an analyst might ask about an agent. In particular, we show that each question cannot be addressed by pure observation alone, but instead requires conducting experiments with systematically chosen manipulations so as to generate the correct causal evidence.

Keywords

Cite

@article{arxiv.2103.03938,
  title  = {Causal Analysis of Agent Behavior for AI Safety},
  author = {Grégoire Déletang and Jordi Grau-Moya and Miljan Martic and Tim Genewein and Tom McGrath and Vladimir Mikulik and Markus Kunesch and Shane Legg and Pedro A. Ortega},
  journal= {arXiv preprint arXiv:2103.03938},
  year   = {2021}
}

Comments

16 pages, 16 figures, 6 tables

R2 v1 2026-06-23T23:49:20.269Z