English
Related papers

Related papers: Amuse: Human-AI Collaborative Songwriting with Mul…

200 papers

Automatic melody generation for pop music has been a long-time aspiration for both AI researchers and musicians. However, learning to generate euphonious melody has turned out to be highly challenging due to a number of factors.…

Lyric-to-melody generation, which generates melody according to given lyrics, is one of the most important automatic music composition tasks. With the rapid development of deep learning, previous works address this task with end-to-end…

Sound · Computer Science 2022-07-13 Chen Zhang , Luchin Chang , Songruoyao Wu , Xu Tan , Tao Qin , Tie-Yan Liu , Kejun Zhang

Word embedding has become an essential means for text-based information retrieval. Typically, word embeddings are learned from large quantities of general and unstructured text data. However, in the domain of music, the word embedding may…

Sound · Computer Science 2024-04-24 SeungHeon Doh , Jongpil Lee , Dasaem Jeong , Juhan Nam

We consider and propose a new problem of retrieving audio files relevant to multimodal design document inputs comprising both textual elements and visual imagery, e.g., birthday/greeting cards. In addition to enhancing user experience,…

Multimedia · Computer Science 2023-03-01 Prachi Singh , Srikrishna Karanam , Sumit Shekhar

Singing-driven 3D head animation is a challenging yet promising task with applications in virtual avatars, entertainment, and education. Unlike speech, singing involves richer emotional nuance, dynamic prosody, and lyric-based semantics,…

Graphics · Computer Science 2025-09-03 Zikai Huang , Yihan Zhou , Xuemiao Xu , Cheng Xu , Xiaofen Xing , Jing Qin , Shengfeng He

Online AI platforms for creating music from text prompts (AI music), such as Suno and Udio, are now being used by hundreds of thousands of users. Some AI music is appearing in advertising, and even charting, in multiple countries. How are…

Information Retrieval · Computer Science 2025-09-16 Luca Casini , Laura Cros Vila , David Dalmazzo , Anna-Kaisa Kaila , Bob L. T. Sturm

Effective human-AI coordination requires artificial agents capable of exhibiting and responding to human-like behaviors while adapting to changing contexts. Imitation learning has emerged as one of the prominent approaches to build such…

Artificial Intelligence · Computer Science 2026-02-25 Rakshit Trivedi , Kartik Sharma , David C Parkes

Automatic Singing Assessment and Singing Information Processing have evolved over the past three decades to support singing pedagogy, performance analysis, and vocal training. While the first approach objectively evaluates a singer's…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Arthur N. dos Santos , Bruno S. Masiero

High-quality, large-scale audio captioning is crucial for advancing audio understanding, yet current automated methods often generate captions that lack fine-grained detail and contextual accuracy, primarily due to their reliance on limited…

Sound · Computer Science 2025-06-03 Shunian Chen , Xinyuan Xie , Zheshu Chen , Liyan Zhao , Owen Lee , Zhan Su , Qilin Sun , Benyou Wang

In pop music, accompaniments are usually played by multiple instruments (tracks) such as drum, bass, string and guitar, and can make a song more expressive and contagious by arranging together with its melody. Previous works usually…

Sound · Computer Science 2020-08-19 Yi Ren , Jinzheng He , Xu Tan , Tao Qin , Zhou Zhao , Tie-Yan Liu

When songs are composed or performed, there is often an intent by the singer/songwriter of expressing feelings or emotions through it. For humans, matching the emotiveness in a musical composition or performance with the subjective…

Human machine interaction is a huge source of inspiration in today's media art and digital design, as machines and humans merge together more and more. Its place in art reflects its growing applications in industry, such as robotics.…

Neural and Evolutionary Computing · Computer Science 2025-07-22 Jules Lecomte , Konrad Zinner , Michael Neumeier , Axel von Arnim

Journaling has long been recognized for fostering emotional awareness and self-reflection, and recent advancements in generative AI offer new opportunities to create personalized music that can enhance these practices. In this study, we…

Human-Computer Interaction · Computer Science 2025-06-03 Joonyoung Park , Hyewon Cho , Hyehyun Chu , Yeeun Lee , Hajin Lim

Voice assistants provide users a new way of interacting with digital products, allowing them to retrieve information and complete tasks with an increased sense of control and flexibility. Such products are comprised of several machine…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-27 Shachaf Poran , Gil Amsalem , Amit Beka , Dmitri Goldenberg

Sentiment analysis, mostly based on text, has been rapidly developing in the last decade and has attracted widespread attention in both academia and industry. However, the information in the real world usually comes from multiple…

Computation and Language · Computer Science 2019-12-12 Feiyang Chen , Ziqian Luo , Yanyan Xu , Dengfeng Ke

Based on recent advances in realistic language modeling (GPT-3) and cross-modal representations (CLIP), Gaud\'i was developed to help designers search for inspirational images using natural language. In the early stages of the design…

Artificial Intelligence · Computer Science 2021-12-09 Victor S. Bursztyn , Jennifer Healey , Vishwa Vinay

We present Hookpad Aria, a generative AI system designed to assist musicians in writing Western pop songs. Our system is seamlessly integrated into Hookpad, a web-based editor designed for the composition of lead sheets: symbolic music…

Sound · Computer Science 2025-02-13 Chris Donahue , Shih-Lun Wu , Yewon Kim , Dave Carlton , Ryan Miyakawa , John Thickstun

The AI community has embraced multi-sensory or multi-modal approaches to advance this generation of AI models to resemble expected intelligent understanding. Combining language and imagery represents a familiar method for specific tasks…

Computation and Language · Computer Science 2023-04-06 David Noever , Samantha Elizabeth Miller Noever

Retrieval-augmented generation can improve audio captioning by incorporating relevant audio-text pairs from a knowledge base. Existing methods typically rely solely on the input audio as a unimodal retrieval query. In contrast, we propose…

Sound · Computer Science 2025-06-11 Choi Changin , Lim Sungjun , Rhee Wonjong

The use of artificial intelligence (AI) to support creative writing has bloomed in recent years. However, it is less well understood how AI compares to on-demand human support. We explored how writers interact with both AI and crowd worker…

Human-Computer Interaction · Computer Science 2024-10-22 Chieh-Yang Huang , Sanjana Gautam , Shannon McClellan Brooks , Ya-Fang Lin , Tiffany Knearem , Ting-Hao 'Kenneth' Huang