English

A Multiscale Visualization of Attention in the Transformer Model

Human-Computer Interaction 2019-06-14 v1 Computation and Language Machine Learning

Abstract

The Transformer is a sequence model that forgoes traditional recurrent architectures in favor of a fully attention-based approach. Besides improving performance, an advantage of using attention is that it can also help to interpret a model by showing how the model assigns weight to different input elements. However, the multi-layer, multi-head attention mechanism in the Transformer model can be difficult to decipher. To make the model more accessible, we introduce an open-source tool that visualizes attention at multiple scales, each of which provides a unique perspective on the attention mechanism. We demonstrate the tool on BERT and OpenAI GPT-2 and present three example use cases: detecting model bias, locating relevant attention heads, and linking neurons to model behavior.

Keywords

Cite

@article{arxiv.1906.05714,
  title  = {A Multiscale Visualization of Attention in the Transformer Model},
  author = {Jesse Vig},
  journal= {arXiv preprint arXiv:1906.05714},
  year   = {2019}
}

Comments

To appear in ACL 2019 (System Demonstrations). arXiv admin note: substantial text overlap with arXiv:1904.02679

R2 v1 2026-06-23T09:52:49.708Z