Discovering Universal Geometry in Embeddings with ICA
Computation and Language
2023-11-03 v2
Abstract
This study utilizes Independent Component Analysis (ICA) to unveil a consistent semantic structure within embeddings of words or images. Our approach extracts independent semantic components from the embeddings of a pre-trained model by leveraging anisotropic information that remains after the whitening process in Principal Component Analysis (PCA). We demonstrate that each embedding can be expressed as a composition of a few intrinsic interpretable axes and that these semantic axes remain consistent across different languages, algorithms, and modalities. The discovery of a universal semantic structure in the geometric patterns of embeddings enhances our understanding of the representations in embeddings.
Cite
@article{arxiv.2305.13175,
title = {Discovering Universal Geometry in Embeddings with ICA},
author = {Hiroaki Yamagiwa and Momose Oyama and Hidetoshi Shimodaira},
journal= {arXiv preprint arXiv:2305.13175},
year = {2023}
}
Comments
29 pages, EMNLP 2023