English

ULTra: Unveiling Latent Token Interpretability in Transformer-Based Understanding and Segmentation

Computer Vision and Pattern Recognition 2026-01-06 v3 Artificial Intelligence Machine Learning

Abstract

Transformers have revolutionized Computer Vision (CV) through self-attention mechanisms. However, their complexity makes latent token representations difficult to interpret. We introduce ULTra, a framework for interpreting Transformer embeddings and uncovering meaningful semantic patterns within them. ULTra enables unsupervised semantic segmentation using pre-trained models without requiring fine-tuning. Additionally, we propose a self-supervised training approach that refines segmentation performance by learning an external transformation matrix without modifying the underlying model. Our method achieves state-of-the-art performance in unsupervised semantic segmentation, outperforming existing segmentation methods. Furthermore, we validate ULTra for model interpretation on both synthetic and real-world scenarios, including Object Selection and interpretable text summarization using LLMs, demonstrating its broad applicability in explaining the semantic structure of latent token representations.

Keywords

Cite

@article{arxiv.2411.12589,
  title  = {ULTra: Unveiling Latent Token Interpretability in Transformer-Based Understanding and Segmentation},
  author = {Hesam Hosseini and Ghazal Hosseini Mighan and Amirabbas Afzali and Sajjad Amini and Amir Houmansadr},
  journal= {arXiv preprint arXiv:2411.12589},
  year   = {2026}
}
R2 v1 2026-06-28T20:05:10.223Z