English

Words that make SENSE: Sensorimotor Norms in Learned Lexical Token Representations

Computation and Language 2026-04-24 v2 Artificial Intelligence

Abstract

While word embeddings derive meaning from co-occurrence patterns, human language understanding is grounded in sensory and motor experience. We present SENSE\text{SENSE} (Sensorimotor (\textbf{S}\text{ensorimotor } Embedding \textbf{E}\text{mbedding } Norm \textbf{N}\text{orm } Scoring \textbf{S}\text{coring } Engine)\textbf{E}\text{ngine}), a learned projection model that predicts Lancaster sensorimotor norms from word lexical embeddings. We also conducted a behavioral study where 281 participants selected which among candidate nonce words evoked specific sensorimotor associations, finding statistically significant correlations between human selection rates and SENSE\text{SENSE} ratings across 6 of the 11 modalities. Sublexical analysis of these nonce words selection rates revealed systematic phonosthemic patterns for the interoceptive norm, suggesting a path towards computationally proposing candidate phonosthemes from text data.

Keywords

Cite

@article{arxiv.2602.00469,
  title  = {Words that make SENSE: Sensorimotor Norms in Learned Lexical Token Representations},
  author = {Abhinav Gupta and Toben H. Mintz and Jesse Thomason},
  journal= {arXiv preprint arXiv:2602.00469},
  year   = {2026}
}

Comments

5 pages, 2 figures, codebase can be found at: https://github.com/abhinav-usc/SENSE-model/tree/main

R2 v1 2026-07-01T09:28:59.243Z