English

Linear Latent World Models in Simple Transformers: A Case Study on Othello-GPT

Machine Learning 2023-10-24 v2 Artificial Intelligence

Abstract

Foundation models exhibit significant capabilities in decision-making and logical deductions. Nonetheless, a continuing discourse persists regarding their genuine understanding of the world as opposed to mere stochastic mimicry. This paper meticulously examines a simple transformer trained for Othello, extending prior research to enhance comprehension of the emergent world model of Othello-GPT. The investigation reveals that Othello-GPT encapsulates a linear representation of opposing pieces, a factor that causally steers its decision-making process. This paper further elucidates the interplay between the linear world representation and causal decision-making, and their dependence on layer depth and model complexity. We have made the code public.

Keywords

Cite

@article{arxiv.2310.07582,
  title  = {Linear Latent World Models in Simple Transformers: A Case Study on Othello-GPT},
  author = {Dean S. Hazineh and Zechen Zhang and Jeffery Chiu},
  journal= {arXiv preprint arXiv:2310.07582},
  year   = {2023}
}
R2 v1 2026-06-28T12:47:30.721Z