English
Related papers

Related papers: Amuse: Human-AI Collaborative Songwriting with Mul…

200 papers

Recent advances in large language models have facilitated the development of unified speech language models (SLMs) capable of supporting multiple speech tasks within a shared architecture. However, tasks such as automatic speech recognition…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-24 Yuke Si , Runyan Yang , Yingying Gao , Junlan Feng , Chao Deng , Shilei Zhang

Analyzing individual emotions during group conversation is crucial in developing intelligent agents capable of natural human-machine interaction. While reliable emotion recognition techniques depend on different modalities (text, audio,…

Content creators often use music to enhance their stories, as it can be a powerful tool to convey emotion. In this paper, our goal is to help creators find music to match the emotion of their story. We focus on text-based stories that can…

Information Retrieval · Computer Science 2021-11-29 Minz Won , Justin Salamon , Nicholas J. Bryan , Gautham J. Mysore , Xavier Serra

This study addresses the deficiency in conventional music recommendation systems by focusing on the vital role of emotions in shaping users music choices. These systems often disregard the emotional context, relying predominantly on past…

Information Retrieval · Computer Science 2023-11-21 Tina Babu , Rekha R Nair , Geetha A

Most musical programming languages are developed purely for coding virtual instruments or algorithmic compositions. Although there has been some work in the domain of musical query languages for music information retrieval, there has been…

Sound · Computer Science 2017-09-08 Donya Quick , Clayton T. Morrison

In this paper, we discuss the conceptualisation and design of embodied AI within an inclusive music-making project. The central case study is Jess+ an intelligent digital score system for shared creativity with a mixed ensemble of…

Human-Computer Interaction · Computer Science 2024-12-10 Craig Vear , Johann Benerradi

The advancement of Machine learning (ML), Large Audio Language Models (LALMs), and autonomous AI agents in Music Information Retrieval (MIR) necessitates a shift from static tagging to rich, human-aligned representation learning. However,…

Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand such signals in human cultural contexts, we introduce AVMeme…

The field of AI-assisted music creation has made significant strides, yet existing systems often struggle to meet the demands of iterative and nuanced music production. These challenges include providing sufficient control over the…

Sound · Computer Science 2024-11-22 Yixiao Zhang

Multimodal models are critical for music understanding tasks, as they capture the complex interplay between audio and lyrics. However, as these models become more prevalent, the need for explainability grows-understanding how these systems…

Music generation has advanced markedly through multimodal deep learning, enabling models to synthesize audio from text and, more recently, from images. However, existing image-conditioned systems suffer from two fundamental limitations: (i)…

Computer Vision and Pattern Recognition · Computer Science 2026-02-20 Ivan Rinaldi , Matteo Mendula , Nicola Fanelli , Florence Levé , Matteo Testi , Giovanna Castellano , Gennaro Vessio

We discuss a novel task, Chorus Recognition, which could potentially benefit downstream tasks such as song search and music summarization. Different from the existing tasks such as music summarization or lyrics summarization relying on…

Information Retrieval · Computer Science 2021-07-01 Jiaan Wang , Zhixu Li , Binbin Gu , Tingyi Zhang , Qingsheng Liu , Zhigang Chen

Towards improving the performance in various music information processing tasks, recent studies exploit different modalities able to capture diverse aspects of music. Such modalities include audio recordings, symbolic music scores,…

Multimedia · Computer Science 2019-02-15 Federico Simonetta , Stavros Ntalampiras , Federico Avanzini

The Music Emotion Recognition (MER) field has seen steady developments in recent years, with contributions from feature engineering, machine learning, and deep learning. The landscape has also shifted from audio-centric systems to bimodal…

Rapid advancements in artificial intelligence have significantly enhanced generative tasks involving music and images, employing both unimodal and multimodal approaches. This research develops a model capable of generating music that…

Sound · Computer Science 2024-09-13 Tanisha Hisariya , Huan Zhang , Jinhua Liang

Nonverbal communication is integral to human interaction, with gestures, facial expressions, and body language conveying critical aspects of intent and emotion. However, existing large language models (LLMs) fail to effectively incorporate…

Artificial Intelligence · Computer Science 2025-06-03 Youngmin Kim , Jiwan Chung , Jisoo Kim , Sunghyun Lee , Sangkyu Lee , Junhyeok Kim , Cheoljong Yang , Youngjae Yu

The present study proposes a novel approach to dream recording by combining non-invasive brain-machine interfaces (BMI), thought-typing software, and generative AI-assisted multimodal software. This method aims to sublimate conscious…

Human-Computer Interaction · Computer Science 2023-04-21 Todd Kelsey

Collaboration is built on trust, and establishing trust with a creative Artificial Intelligence is difficult when the decision process or internal state driving its behaviour isn't exposed. When human musicians improvise together, a number…

Human-Computer Interaction · Computer Science 2019-02-19 Jon McCormack , Toby Gifford , Patrick Hutchings , Maria Teresa Llano Rodriguez , Matthew Yee-King , Mark d'Inverno

Recent advancements in music large language models (LLMs) have significantly improved music understanding tasks, which involve the model's ability to analyze and interpret various musical elements. These improvements primarily focused on…

Sound · Computer Science 2025-09-24 Zhuoyuan Mao , Mengjie Zhao , Qiyu Wu , Hiromi Wakaki , Yuki Mitsufuji

In the domain of multimodal intent recognition (MIR), the objective is to recognize human intent by integrating a variety of modalities, such as language text, body gestures, and tones. However, existing approaches face difficulties…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Yaomin Shen , Xiaojian Lin , Wei Fan