English

Enhancing Video Music Recommendation with Transformer-Driven Audio-Visual Embeddings

Multimedia 2025-03-10 v1

Abstract

A fitting soundtrack can help a video better convey its content and provide a better immersive experience. This paper introduces a novel approach utilizing self-supervised learning and contrastive learning to automatically recommend audio for video content, thereby eliminating the need for manual labeling. We use a dual-branch cross-modal embedding model that maps both audio and video features into a common low-dimensional space. The fit of various audio-video pairs can then be mod-eled as inverse distance measure. In addition, a comparative analysis of various temporal encoding methods is presented, emphasizing the effectiveness of transformers in managing the temporal information of audio-video matching tasks. Through multiple experiments, we demonstrate that our model TIVM, which integrates transformer encoders and using InfoN Celoss, significantly improves the performance of audio-video matching and surpasses traditional methods.

Keywords

Cite

@article{arxiv.2503.05008,
  title  = {Enhancing Video Music Recommendation with Transformer-Driven Audio-Visual Embeddings},
  author = {Shimiao Liu and Alexander Lerch},
  journal= {arXiv preprint arXiv:2503.05008},
  year   = {2025}
}

Comments

2024 IEEE 5th International Symposium on the Internet of Sounds (IS2), Erlangen, Germany, 2024, pp. 1-6

R2 v1 2026-06-28T22:10:05.903Z