English

Learning Embodied Semantics via Music and Dance Semiotic Correlations

Computer Vision and Pattern Recognition 2021-12-13 v1 Machine Learning Sound Audio and Speech Processing

Abstract

Music semantics is embodied, in the sense that meaning is biologically mediated by and grounded in the human body and brain. This embodied cognition perspective also explains why music structures modulate kinetic and somatosensory perception. We leverage this aspect of cognition, by considering dance as a proxy for music perception, in a statistical computational model that learns semiotic correlations between music audio and dance video. We evaluate the ability of this model to effectively capture underlying semantics in a cross-modal retrieval task. Quantitative results, validated with statistical significance testing, strengthen the body of evidence for embodied cognition in music and show the model can recommend music audio for dance video queries and vice-versa.

Keywords

Cite

@article{arxiv.1903.10534,
  title  = {Learning Embodied Semantics via Music and Dance Semiotic Correlations},
  author = {Francisco Afonso Raposo and David Martins de Matos and Ricardo Ribeiro},
  journal= {arXiv preprint arXiv:1903.10534},
  year   = {2021}
}

Comments

24 pages, 1 figure, 5 tables

R2 v1 2026-06-23T08:18:41.125Z