English

Cross-modal Embeddings for Video and Audio Retrieval

Information Retrieval 2018-01-09 v1 Computer Vision and Pattern Recognition Sound Audio and Speech Processing

Abstract

The increasing amount of online videos brings several opportunities for training self-supervised neural networks. The creation of large scale datasets of videos such as the YouTube-8M allows us to deal with this large amount of data in manageable way. In this work, we find new ways of exploiting this dataset by taking advantage of the multi-modal information it provides. By means of a neural network, we are able to create links between audio and visual documents, by projecting them into a common region of the feature space, obtaining joint audio-visual embeddings. These links are used to retrieve audio samples that fit well to a given silent video, and also to retrieve images that match a given a query audio. The results in terms of Recall@K obtained over a subset of YouTube-8M videos show the potential of this unsupervised approach for cross-modal feature learning. We train embeddings for both scales and assess their quality in a retrieval problem, formulated as using the feature extracted from one modality to retrieve the most similar videos based on the features computed in the other modality.

Keywords

Cite

@article{arxiv.1801.02200,
  title  = {Cross-modal Embeddings for Video and Audio Retrieval},
  author = {Didac Surís and Amanda Duarte and Amaia Salvador and Jordi Torres and Xavier Giró-i-Nieto},
  journal= {arXiv preprint arXiv:1801.02200},
  year   = {2018}
}

Comments

6 pages, 3 figures

R2 v1 2026-06-22T23:38:36.414Z