English

Flexible Control in Symbolic Music Generation via Musical Metadata

Sound 2024-09-13 v1 Multimedia Audio and Speech Processing

Abstract

In this work, we introduce the demonstration of symbolic music generation, focusing on providing short musical motifs that serve as the central theme of the narrative. For the generation, we adopt an autoregressive model which takes musical metadata as inputs and generates 4 bars of multitrack MIDI sequences. During training, we randomly drop tokens from the musical metadata to guarantee flexible control. It provides users with the freedom to select input types while maintaining generative performance, enabling greater flexibility in music composition. We validate the effectiveness of the strategy through experiments in terms of model capacity, musical fidelity, diversity, and controllability. Additionally, we scale up the model and compare it with other music generation model through a subjective test. Our results indicate its superiority in both control and music quality. We provide a URL link https://www.youtube.com/watch?v=-0drPrFJdMQ to our demonstration video.

Keywords

Cite

@article{arxiv.2409.07467,
  title  = {Flexible Control in Symbolic Music Generation via Musical Metadata},
  author = {Sangjun Han and Jiwon Ham and Chaeeun Lee and Heejin Kim and Soojong Do and Sihyuk Yi and Jun Seo and Seoyoon Kim and Yountae Jung and Woohyung Lim},
  journal= {arXiv preprint arXiv:2409.07467},
  year   = {2024}
}