Content-based Music Similarity with Triplet Networks
Machine Learning
2022-12-08 v2 Sound
Audio and Speech Processing
Abstract
We explore the feasibility of using triplet neural networks to embed songs based on content-based music similarity. Our network is trained using triplets of songs such that two songs by the same artist are embedded closer to one another than to a third song by a different artist. We compare two models that are trained using different ways of picking this third song: at random vs. based on shared genre labels. Our experiments are conducted using songs from the Free Music Archive and use standard audio features. The initial results show that shallow Siamese networks can be used to embed music for a simple artist retrieval task.
Keywords
Cite
@article{arxiv.2008.04938,
title = {Content-based Music Similarity with Triplet Networks},
author = {Joseph Cleveland and Derek Cheng and Michael Zhou and Thorsten Joachims and Douglas Turnbull},
journal= {arXiv preprint arXiv:2008.04938},
year = {2022}
}