English

Multimodal Dataset Normalization and Perceptual Validation for Music-Taste Correspondences

Sound 2026-04-14 v1 Machine Learning Multimedia Audio and Speech Processing

Abstract

Collecting large, aligned cross-modal datasets for music-flavor research is difficult because perceptual experiments are costly and small by design. We address this bottleneck through two complementary experiments. The first tests whether audio-flavor correlations, feature-importance rankings, and latent-factor structure transfer from an experimental soundtracks collection (257~tracks with human annotations) to a large FMA-derived corpus (\sim49,300 segments with synthetic labels). The second validates computational flavor targets -- derived from food chemistry via a reproducible pipeline -- against human perception in an online listener study (49~participants, 20~tracks). Results from both experiments converge: the quantitative transfer analysis confirms that cross-modal structure is preserved across supervision regimes, and the perceptual evaluation shows significant alignment between computational targets and listener ratings (permutation p<0.0001p<0.0001, Mantel r=0.45r=0.45, Procrustes m2=0.51m^2=0.51). Together, these findings support the conclusion that sonic seasoning effects are present in synthetic FMA annotations. We release datasets and companion code to support reproducible cross-modal AI research.

Keywords

Cite

@article{arxiv.2604.10632,
  title  = {Multimodal Dataset Normalization and Perceptual Validation for Music-Taste Correspondences},
  author = {Matteo Spanio and Valentina Frezzato and Antonio Rodà},
  journal= {arXiv preprint arXiv:2604.10632},
  year   = {2026}
}

Comments

Submitted to SMC2026

R2 v1 2026-07-01T12:05:00.838Z