English

Foundation Models for Music: A Survey

Sound 2024-09-04 v3 Artificial Intelligence Computation and Language Machine Learning Audio and Speech Processing

Abstract

In recent years, foundation models (FMs) such as large language models (LLMs) and latent diffusion models (LDMs) have profoundly impacted diverse sectors, including music. This comprehensive review examines state-of-the-art (SOTA) pre-trained models and foundation models in music, spanning from representation learning, generative learning and multimodal learning. We first contextualise the significance of music in various industries and trace the evolution of AI in music. By delineating the modalities targeted by foundation models, we discover many of the music representations are underexplored in FM development. Then, emphasis is placed on the lack of versatility of previous methods on diverse music applications, along with the potential of FMs in music understanding, generation and medical application. By comprehensively exploring the details of the model pre-training paradigm, architectural choices, tokenisation, finetuning methodologies and controllability, we emphasise the important topics that should have been well explored, like instruction tuning and in-context learning, scaling law and emergent ability, as well as long-sequence modelling etc. A dedicated section presents insights into music agents, accompanied by a thorough analysis of datasets and evaluations essential for pre-training and downstream tasks. Finally, by underscoring the vital importance of ethical considerations, we advocate that following research on FM for music should focus more on such issues as interpretability, transparency, human responsibility, and copyright issues. The paper offers insights into future challenges and trends on FMs for music, aiming to shape the trajectory of human-AI collaboration in the music realm.

Keywords

Cite

@article{arxiv.2408.14340,
  title  = {Foundation Models for Music: A Survey},
  author = {Yinghao Ma and Anders Øland and Anton Ragni and Bleiz MacSen Del Sette and Charalampos Saitis and Chris Donahue and Chenghua Lin and Christos Plachouras and Emmanouil Benetos and Elona Shatri and Fabio Morreale and Ge Zhang and György Fazekas and Gus Xia and Huan Zhang and Ilaria Manco and Jiawen Huang and Julien Guinot and Liwei Lin and Luca Marinelli and Max W. Y. Lam and Megha Sharma and Qiuqiang Kong and Roger B. Dannenberg and Ruibin Yuan and Shangda Wu and Shih-Lun Wu and Shuqi Dai and Shun Lei and Shiyin Kang and Simon Dixon and Wenhu Chen and Wenhao Huang and Xingjian Du and Xingwei Qu and Xu Tan and Yizhi Li and Zeyue Tian and Zhiyong Wu and Zhizheng Wu and Ziyang Ma and Ziyu Wang},
  journal= {arXiv preprint arXiv:2408.14340},
  year   = {2024}
}
R2 v1 2026-06-28T18:24:05.385Z