Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset
Abstract
Recent years have seen many audio-domain text-to-music generation models that rely on large amounts of text-audio pairs for training. However, symbolic-domain controllable music generation has lagged behind partly due to the lack of a large-scale symbolic music dataset with extensive metadata and captions. In this work, we present MetaScore, a new dataset consisting of 963K musical scores paired with rich metadata, including free-form user-annotated tags, collected from an online music forum. To approach text-to-music generation, We employ a pretrained large language model (LLM) to generate pseudo-natural language captions for music from its metadata tags. With the LLM-enhanced MetaScore, we train a text-conditioned music generation model that learns to generate symbolic music from the pseudo captions, allowing control of instruments, genre, composer, complexity and other free-form music descriptors. In addition, we train a tag-conditioned system that supports a predefined set of tags available in MetaScore. Our experimental results show that both the proposed text-to-music and tags-to-music models outperform a baseline text-to-music model in a listening test. While a concurrent work Text2MIDI also supports free-form text input, our models achieve comparable performance. Moreover, the text-to-music system offers a more natural interface than the tags-to-music model, as it allows users to provide free-form natural language prompts.
Cite
@article{arxiv.2410.02084,
title = {Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset},
author = {Weihan Xu and Julian McAuley and Taylor Berg-Kirkpatrick and Shlomo Dubnov and Hao-Wen Dong},
journal= {arXiv preprint arXiv:2410.02084},
year = {2025}
}
Comments
Accepted at ISMIR 2025