This paper introduces effective design choices for text-to-music retrieval systems. An ideal text-based retrieval system would support various input queries such as pre-defined tags, unseen tags, and sentence-level descriptions. In reality, most previous works mainly focused on a single query type (tag or sentence) which may not generalize to another input type. Hence, we review recent text-based music retrieval systems using our proposed benchmark in two main aspects: input text representation and training objectives. Our findings enable a universal text-to-music retrieval system that achieves comparable retrieval performances in both tag- and sentence-level inputs. Furthermore, the proposed multimodal representation generalizes to 9 different downstream music classification tasks. We present the code and demo online.
@article{arxiv.2211.14558,
title = {Toward Universal Text-to-Music Retrieval},
author = {SeungHeon Doh and Minz Won and Keunwoo Choi and Juhan Nam},
journal= {arXiv preprint arXiv:2211.14558},
year = {2022}
}