English

Integrating Text-to-Music Models with Language Models: Composing Long Structured Music Pieces

Sound 2024-10-08 v3 Machine Learning Audio and Speech Processing

Abstract

Recent music generation methods based on transformers have a context window of up to a minute. The music generated by these methods is largely unstructured beyond the context window. With a longer context window, learning long-scale structures from musical data is a prohibitively challenging problem. This paper proposes integrating a text-to-music model with a large language model to generate music with form. The papers discusses the solutions to the challenges of such integration. The experimental results show that the proposed method can generate 2.5-minute-long music that is highly structured, strongly organized, and cohesive.

Keywords

Cite

@article{arxiv.2410.00344,
  title  = {Integrating Text-to-Music Models with Language Models: Composing Long Structured Music Pieces},
  author = {Lilac Atassi},
  journal= {arXiv preprint arXiv:2410.00344},
  year   = {2024}
}

Comments

arXiv admin note: substantial text overlap with arXiv:2404.11976

R2 v1 2026-06-28T19:03:17.416Z