English

Bridging Compositional and Distributional Semantics: A Survey on Latent Semantic Geometry via AutoEncoder

Computation and Language 2026-04-16 v4

Abstract

Integrating compositional and symbolic properties into current distributional semantic spaces can enhance the interpretability, controllability, compositionality, and generalisation capabilities of Transformer-based auto-regressive language models (LMs). In this survey, we offer a novel perspective on latent space geometry through the lens of compositional semantics, a direction we refer to as \textit{semantic representation learning}. This direction enables a bridge between symbolic and distributional semantics, helping to mitigate the gap between them. We review and compare three mainstream autoencoder architectures-Variational AutoEncoder (VAE), Vector Quantised VAE (VQVAE), and Sparse AutoEncoder (SAE)-and examine the distinctive latent geometries they induce in relation to semantic structure and interpretability.

Keywords

Cite

@article{arxiv.2506.20083,
  title  = {Bridging Compositional and Distributional Semantics: A Survey on Latent Semantic Geometry via AutoEncoder},
  author = {Yingji Zhang and Danilo S. Carvalho and André Freitas},
  journal= {arXiv preprint arXiv:2506.20083},
  year   = {2026}
}

Comments

In progress

R2 v1 2026-07-01T03:32:26.614Z