English
Related papers

Related papers: MusicLM: Generating Music From Text

200 papers

Loopable music generation systems enable diverse applications, but they often lack controllability and customization capabilities. We argue that enhancing controllability can enrich these models, with emotional expression being a crucial…

Sound · Computer Science 2024-01-26 Wenqian Cui , Pedro Sarmento , Mathieu Barthet

Text-to-music models have revolutionized the creative landscape, offering new possibilities for music creation. Yet their integration into musicians workflows remains underexplored. This paper presents a case study on how TTM models impact…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 Francesca Ronchini , Luca Comanducci , Simone Marcucci , Fabio Antonacci

Generative models guided by text prompts are increasingly becoming more popular. However, no text-to-MIDI models currently exist due to the lack of a captioned MIDI dataset. This work aims to enable research that combines LLMs with symbolic…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-08 Jan Melechovsky , Abhinaba Roy , Dorien Herremans

Large Language Models (LLMs) show promise in lyric-to-melody generation, but models trained with Supervised Fine-Tuning (SFT) often produce musically implausible melodies with issues like poor rhythm and unsuitable vocal ranges, a…

Sound · Computer Science 2026-04-21 Hao Meng , Siyuan Zheng , Shuran Zhou , Qiangqiang Wang , Yang Song

Music generated by deep learning methods often suffers from a lack of coherence and long-term organization. Yet, multi-scale hierarchical structure is a distinctive feature of music signals. To leverage this information, we propose a…

Sound · Computer Science 2024-02-29 Manvi Agarwal , Changhong Wang , Gaël Richard

Research on large language models has advanced significantly across text, speech, images, and videos. However, multi-modal music understanding and generation remain underexplored due to the lack of well-annotated datasets. To address this,…

Sound · Computer Science 2024-12-10 Shansong Liu , Atin Sakkeer Hussain , Qilong Wu , Chenshuo Sun , Ying Shan

Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-ZERO, a video-to-music generation approach that generates time-aligned…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Yan-Bo Lin , Jonah Casebeer , Long Mai , Aniruddha Mahapatra , Gedas Bertasius , Nicholas J. Bryan

Gesture-driven music generation is an emerging human-computer interaction paradigm for touch-free and expressive musical interaction. However, many existing approaches treat the task as isolated gesture classification or map gestures to…

Multimedia · Computer Science 2026-04-29 Rathinaraja Jeyaraj , Barathi Subramanian , Kapilya Gangadharan , Anand Paul

We introduce the text-to-instrument task, which aims at generating sample-based musical instruments based on textual prompts. Accordingly, we propose InstrumentGen, a model that extends a text-prompted generative audio framework to…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-09 Shahan Nercessian , Johannes Imort

Automatic song writing is a topic of significant practical interest. However, its research is largely hindered by the lack of training data due to copyright concerns and challenged by its creative nature. Most noticeably, prior works often…

While most music generation models generate a mixture of stems (in mono or stereo), we propose to train a multi-stem generative model with 3 stems (bass, drums and other) that learn the musical dependencies between them. To do so, we train…

Sound · Computer Science 2025-01-08 Simon Rouard , Robin San Roman , Yossi Adi , Axel Roebel

Music is a universal language that can communicate emotions and feelings. It forms an essential part of the whole spectrum of creative media, ranging from movies to social media posts. Machine learning models that can synthesize music are…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Sanjoy Chowdhury , Sayan Nag , K J Joseph , Balaji Vasan Srinivasan , Dinesh Manocha

Musical mode is one of the most critical element that establishes the framework of pitch organization and determines the harmonic relationships. Previous works often use the simplistic and rigid alignment method, and overlook the diversity…

Sound · Computer Science 2025-01-15 Qian Liang , Yi Zeng , Menghaoran Tang

While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind. In this paper, we bridge this critical gap by establishing a comprehensive…

We introduce a film score generation framework to harmonize visual pixels and music melodies utilizing a latent diffusion model. Our framework processes film clips as input and generates music that aligns with a general theme while offering…

Multimedia · Computer Science 2024-11-13 F. Qi , L. Ni , C. Xu

Significant progress has been made in text-to-video generation through the use of powerful generative models and large-scale internet data. However, substantial challenges remain in precisely controlling individual concepts within the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-30 Hanxin Zhu , Tianyu He , Anni Tang , Junliang Guo , Zhibo Chen , Jiang Bian

Despite advances in deep algorithmic music generation, evaluation of generated samples often relies on human evaluation, which is subjective and costly. We focus on designing a homogeneous, objective framework for evaluating samples of…

We present HAFM, a system that generates instrumental music audio to accompany input vocals. Given isolated singing voice, HAFM produces a coherent instrumental accompaniment that can be directly mixed with the input to create complete…

Sound · Computer Science 2026-04-14 Jian Zhu , Jianwei Cui , Shihao Chen , Yubang Zhang , Cheng Luo

We propose Composition Sampling, a simple but effective method to generate diverse outputs for conditional generation of higher quality compared to previous stochastic decoding strategies. It builds on recently proposed plan-based neural…

Computation and Language · Computer Science 2022-03-30 Shashi Narayan , Gonçalo Simões , Yao Zhao , Joshua Maynez , Dipanjan Das , Michael Collins , Mirella Lapata

We present TALKPLAY, a novel multimodal music recommendation system that reformulates recommendation as a token generation problem using large language models (LLMs). By leveraging the instruction-following and natural language generation…

Information Retrieval · Computer Science 2025-05-27 Seungheon Doh , Keunwoo Choi , Juhan Nam