English

Exploring compressibility of transformer based text-to-music (TTM) models

Audio and Speech Processing 2024-06-26 v1 Multimedia Sound

Abstract

State-of-the art Text-To-Music (TTM) generative AI models are large and require desktop or server class compute, making them infeasible for deployment on mobile phones. This paper presents an analysis of trade-offs between model compression and generation performance of TTM models. We study compression through knowledge distillation and specific modifications that enable applicability over the various components of the TTM model (encoder, generative model and the decoder). Leveraging these methods we create TinyTTM (89.2M params) that achieves a FAD of 3.66 and KL of 1.32 on MusicBench dataset, better than MusicGen-Small (557.6M params) but not lower than MusicGen-small fine-tuned on MusicBench.

Keywords

Cite

@article{arxiv.2406.17159,
  title  = {Exploring compressibility of transformer based text-to-music (TTM) models},
  author = {Vasileios Moschopoulos and Thanasis Kotsiopoulos and Pablo Peso Parada and Konstantinos Nikiforidis and Alexandros Stergiadis and Gerasimos Papakostas and Md Asif Jalal and Jisi Zhang and Anastasios Drosou and Karthikeyan Saravanan},
  journal= {arXiv preprint arXiv:2406.17159},
  year   = {2024}
}

Comments

Proceedings of INTERSPEECH 2024