English

The Interpretability of Codebooks in Model-Based Reinforcement Learning is Limited

Artificial Intelligence 2024-07-30 v1 Machine Learning

Abstract

Interpretability of deep reinforcement learning systems could assist operators with understanding how they interact with their environment. Vector quantization methods -- also called codebook methods -- discretize a neural network's latent space that is often suggested to yield emergent interpretability. We investigate whether vector quantization in fact provides interpretability in model-based reinforcement learning. Our experiments, conducted in the reinforcement learning environment Crafter, show that the codes of vector quantization models are inconsistent, have no guarantee of uniqueness, and have a limited impact on concept disentanglement, all of which are necessary traits for interpretability. We share insights on why vector quantization may be fundamentally insufficient for model interpretability.

Keywords

Cite

@article{arxiv.2407.19532,
  title  = {The Interpretability of Codebooks in Model-Based Reinforcement Learning is Limited},
  author = {Kenneth Eaton and Jonathan Balloch and Julia Kim and Mark Riedl},
  journal= {arXiv preprint arXiv:2407.19532},
  year   = {2024}
}
R2 v1 2026-06-28T17:55:57.920Z