中文
相关论文

相关论文: Amuse: Human-AI Collaborative Songwriting with Mul…

200 篇论文

We present TALKPLAY, a novel multimodal music recommendation system that reformulates recommendation as a token generation problem using large language models (LLMs). By leveraging the instruction-following and natural language generation…

信息检索 · 计算机科学 2025-05-27 Seungheon Doh , Keunwoo Choi , Juhan Nam

The rise of "bedroom producers" has democratized music creation, while challenging producers to objectively evaluate their work. To address this, we present AI TrackMate, an LLM-based music chatbot designed to provide constructive feedback…

声音 · 计算机科学 2024-12-10 Yi-Lin Jiang , Chia-Ho Hsiung , Yen-Tung Yeh , Lu-Rong Chen , Bo-Yu Chen

Large Language Models (LLM) have shown encouraging progress in multimodal understanding and generation tasks. However, how to design a human-aligned and interpretable melody composition system is still under-explored. To solve this problem,…

声音 · 计算机科学 2024-03-08 Xia Liang , Xingjian Du , Jiaju Lin , Pei Zou , Yuan Wan , Bilei Zhu

With the rise of artificial intelligence in recent years, there has been a rapid increase in its application towards creative domains, including music. There exist many systems built that apply machine learning approaches to the problem of…

人机交互 · 计算机科学 2025-04-22 Renaud Bougueng Tchemeube , Jeff Ens , Philippe Pasquier

In recent years, remarkable advancements in artificial intelligence-generated content (AIGC) have been achieved in the fields of image synthesis and text generation, generating content comparable to that produced by humans. However, the…

声音 · 计算机科学 2025-01-16 Sida Tian , Can Zhang , Wei Yuan , Wei Tan , Wenjie Zhu

Multimodal models that jointly process audio and language hold great promise in audio understanding and are increasingly being adopted in the music domain. By allowing users to query via text and obtain information about a given audio…

声音 · 计算机科学 2024-08-05 Benno Weck , Ilaria Manco , Emmanouil Benetos , Elio Quinton , George Fazekas , Dmitry Bogdanov

Multi-modal embeddings form the foundation for vision-language models, such as CLIP embeddings, the most widely used text-image embeddings. However, these embeddings are vulnerable to subtle misalignment of cross-modal features, resulting…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Yilin Ye , Shishi Xiao , Xingchen Zeng , Wei Zeng

Music is a potent form of expression that can communicate, accentuate or even create the emotions of an individual or a collective. Both historically and in contemporary experiences, musical expression was and is commonly instrumentalized…

信息检索 · 计算机科学 2024-11-12 Abhishek Kaushik , Kayla Rush

We train a model to generate images from multimodal prompts of interleaved text and images such as "a <picture of a man> man and his <picture of a dog> dog in an <picture of a cartoon> animated style." We bootstrap a multimodal dataset by…

计算机视觉与模式识别 · 计算机科学 2024-09-13 William Berman , Alexander Peysakhovich

Amid the rising intersection of generative AI and human artistic processes, this study probes the critical yet less-explored terrain of alignment in human-centric automatic song composition. We propose a novel task of Colloquial…

声音 · 计算机科学 2024-07-12 Zihao Wang , Haoxuan Liu , Jiaxing Yu , Tao Zhang , Yan Liu , Kejun Zhang

Generative models guided by text prompts are increasingly becoming more popular. However, no text-to-MIDI models currently exist due to the lack of a captioned MIDI dataset. This work aims to enable research that combines LLMs with symbolic…

音频与语音处理 · 电气工程与系统科学 2025-08-08 Jan Melechovsky , Abhinaba Roy , Dorien Herremans

Multimodal scene search of conversations is essential for unlocking valuable insights into social dynamics and enhancing our communication. While experts in conversational analysis have their own knowledge and skills to find key scenes, a…

人机交互 · 计算机科学 2024-02-20 Riku Arakawa , Kiyosu Maeda , Hiromu Yakura

Art has long been a profound medium for expressing emotions. While existing image stylization methods effectively transform visual appearance, they often overlook the emotional impact carried by styles. To bridge this gap, we introduce…

计算机视觉与模式识别 · 计算机科学 2025-12-08 Jingyuan Yang , Zihuan Bai , Hui Huang

In multimodal assistant, where vision is also one of the input modalities, the identification of user intent becomes a challenging task as visual input can influence the outcome. Current digital assistants take spoken input and try to…

计算机视觉与模式识别 · 计算机科学 2021-11-02 Alkesh Patel , Joel Ruben Antony Moniz , Roman Nguyen , Nick Tzou , Hadas Kotek , Vincent Renkens

Machine generation of symbolic music and digital audio are hot topics but there have been relatively few digital musical instruments that integrate generative AI. Present musical AI tools are not artist centred and do not support…

声音 · 计算机科学 2026-04-28 Charles Patrick Martin

Controllable music generation plays a vital role in human-AI music co-creation. While Large Language Models (LLMs) have shown promise in generating high-quality music, their focus on autoregressive generation limits their utility in music…

声音 · 计算机科学 2024-10-08 Liwei Lin , Gus Xia , Yixiao Zhang , Junyan Jiang

A song is a combination of singing voice and accompaniment. However, existing works focus on singing voice synthesis and music generation independently. Little attention was paid to explore song synthesis. In this work, we propose a novel…

音频与语音处理 · 电气工程与系统科学 2024-05-21 Zhiqing Hong , Rongjie Huang , Xize Cheng , Yongqi Wang , Ruiqi Li , Fuming You , Zhou Zhao , Zhimeng Zhang

Jamming requires coordination, anticipation, and collaborative creativity between musicians. Current generative models of music produce expressive output but are not able to generate in an \emph{online} manner, meaning simultaneously with…

The increasing prevalence of AI-generated content alongside human-written text underscores the need for reliable discrimination methods. To address this challenge, we propose a novel framework with textual embeddings from Pre-trained…

计算与语言 · 计算机科学 2024-11-04 Arjun Ramesh Kaushik , Sunil Rufus R P , Nalini Ratha

Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for automatically generating music videos from any…