English
Related papers

Related papers: Amuse: Human-AI Collaborative Songwriting with Mul…

200 papers

We present TALKPLAY, a novel multimodal music recommendation system that reformulates recommendation as a token generation problem using large language models (LLMs). By leveraging the instruction-following and natural language generation…

Information Retrieval · Computer Science 2025-05-27 Seungheon Doh , Keunwoo Choi , Juhan Nam

The rise of "bedroom producers" has democratized music creation, while challenging producers to objectively evaluate their work. To address this, we present AI TrackMate, an LLM-based music chatbot designed to provide constructive feedback…

Sound · Computer Science 2024-12-10 Yi-Lin Jiang , Chia-Ho Hsiung , Yen-Tung Yeh , Lu-Rong Chen , Bo-Yu Chen

Large Language Models (LLM) have shown encouraging progress in multimodal understanding and generation tasks. However, how to design a human-aligned and interpretable melody composition system is still under-explored. To solve this problem,…

Sound · Computer Science 2024-03-08 Xia Liang , Xingjian Du , Jiaju Lin , Pei Zou , Yuan Wan , Bilei Zhu

With the rise of artificial intelligence in recent years, there has been a rapid increase in its application towards creative domains, including music. There exist many systems built that apply machine learning approaches to the problem of…

Human-Computer Interaction · Computer Science 2025-04-22 Renaud Bougueng Tchemeube , Jeff Ens , Philippe Pasquier

In recent years, remarkable advancements in artificial intelligence-generated content (AIGC) have been achieved in the fields of image synthesis and text generation, generating content comparable to that produced by humans. However, the…

Sound · Computer Science 2025-01-16 Sida Tian , Can Zhang , Wei Yuan , Wei Tan , Wenjie Zhu

Multimodal models that jointly process audio and language hold great promise in audio understanding and are increasingly being adopted in the music domain. By allowing users to query via text and obtain information about a given audio…

Sound · Computer Science 2024-08-05 Benno Weck , Ilaria Manco , Emmanouil Benetos , Elio Quinton , George Fazekas , Dmitry Bogdanov

Multi-modal embeddings form the foundation for vision-language models, such as CLIP embeddings, the most widely used text-image embeddings. However, these embeddings are vulnerable to subtle misalignment of cross-modal features, resulting…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Yilin Ye , Shishi Xiao , Xingchen Zeng , Wei Zeng

Music is a potent form of expression that can communicate, accentuate or even create the emotions of an individual or a collective. Both historically and in contemporary experiences, musical expression was and is commonly instrumentalized…

Information Retrieval · Computer Science 2024-11-12 Abhishek Kaushik , Kayla Rush

We train a model to generate images from multimodal prompts of interleaved text and images such as "a <picture of a man> man and his <picture of a dog> dog in an <picture of a cartoon> animated style." We bootstrap a multimodal dataset by…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 William Berman , Alexander Peysakhovich

Amid the rising intersection of generative AI and human artistic processes, this study probes the critical yet less-explored terrain of alignment in human-centric automatic song composition. We propose a novel task of Colloquial…

Sound · Computer Science 2024-07-12 Zihao Wang , Haoxuan Liu , Jiaxing Yu , Tao Zhang , Yan Liu , Kejun Zhang

Generative models guided by text prompts are increasingly becoming more popular. However, no text-to-MIDI models currently exist due to the lack of a captioned MIDI dataset. This work aims to enable research that combines LLMs with symbolic…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-08 Jan Melechovsky , Abhinaba Roy , Dorien Herremans

Multimodal scene search of conversations is essential for unlocking valuable insights into social dynamics and enhancing our communication. While experts in conversational analysis have their own knowledge and skills to find key scenes, a…

Human-Computer Interaction · Computer Science 2024-02-20 Riku Arakawa , Kiyosu Maeda , Hiromu Yakura

Art has long been a profound medium for expressing emotions. While existing image stylization methods effectively transform visual appearance, they often overlook the emotional impact carried by styles. To bridge this gap, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Jingyuan Yang , Zihuan Bai , Hui Huang

In multimodal assistant, where vision is also one of the input modalities, the identification of user intent becomes a challenging task as visual input can influence the outcome. Current digital assistants take spoken input and try to…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Alkesh Patel , Joel Ruben Antony Moniz , Roman Nguyen , Nick Tzou , Hadas Kotek , Vincent Renkens

Machine generation of symbolic music and digital audio are hot topics but there have been relatively few digital musical instruments that integrate generative AI. Present musical AI tools are not artist centred and do not support…

Sound · Computer Science 2026-04-28 Charles Patrick Martin

Controllable music generation plays a vital role in human-AI music co-creation. While Large Language Models (LLMs) have shown promise in generating high-quality music, their focus on autoregressive generation limits their utility in music…

Sound · Computer Science 2024-10-08 Liwei Lin , Gus Xia , Yixiao Zhang , Junyan Jiang

A song is a combination of singing voice and accompaniment. However, existing works focus on singing voice synthesis and music generation independently. Little attention was paid to explore song synthesis. In this work, we propose a novel…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-21 Zhiqing Hong , Rongjie Huang , Xize Cheng , Yongqi Wang , Ruiqi Li , Fuming You , Zhou Zhao , Zhimeng Zhang

Jamming requires coordination, anticipation, and collaborative creativity between musicians. Current generative models of music produce expressive output but are not able to generate in an \emph{online} manner, meaning simultaneously with…

The increasing prevalence of AI-generated content alongside human-written text underscores the need for reliable discrimination methods. To address this challenge, we propose a novel framework with textual embeddings from Pre-trained…

Computation and Language · Computer Science 2024-11-04 Arjun Ramesh Kaushik , Sunil Rufus R P , Nalini Ratha

Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for automatically generating music videos from any…