English

Minimum Levels of Interpretability for Artificial Moral Agents

Artificial Intelligence 2023-07-04 v1 Computers and Society

Abstract

As artificial intelligence (AI) models continue to scale up, they are becoming more capable and integrated into various forms of decision-making systems. For models involved in moral decision-making, also known as artificial moral agents (AMA), interpretability provides a way to trust and understand the agent's internal reasoning mechanisms for effective use and error correction. In this paper, we provide an overview of this rapidly-evolving sub-field of AI interpretability, introduce the concept of the Minimum Level of Interpretability (MLI) and recommend an MLI for various types of agents, to aid their safe deployment in real-world settings.

Keywords

Cite

@article{arxiv.2307.00660,
  title  = {Minimum Levels of Interpretability for Artificial Moral Agents},
  author = {Avish Vijayaraghavan and Cosmin Badea},
  journal= {arXiv preprint arXiv:2307.00660},
  year   = {2023}
}
R2 v1 2026-06-28T11:20:13.077Z