English
Related papers

Related papers: The Interpretation Gap in Text-to-Music Generation…

200 papers

Generative AI has demonstrated impressive performance in various fields, among which speech synthesis is an interesting direction. With the diffusion model as the most popular generative model, numerous works have attempted two active…

In this paper, we consider a novel research problem: music-to-text synaesthesia. Different from the classical music tagging problem that classifies a music recording into pre-defined categories, music-to-text synaesthesia aims to generate…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-09 Zhihuan Kuang , Shi Zong , Jianbing Zhang , Jiajun Chen , Hongfu Liu

Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for automatically generating music videos from any…

This paper is a survey and an analysis of different ways of using deep learning (deep artificial neural networks) to generate musical content. We propose a methodology based on five dimensions for our analysis: Objective - What musical…

Sound · Computer Science 2019-08-09 Jean-Pierre Briot , Gaëtan Hadjeres , François-David Pachet

In the task of generating music, the art factor plays a big role and is a great challenge for AI. Previous work involving adversarial training to produce new music pieces and modeling the compatibility of variety in music (beats, tempo,…

Sound · Computer Science 2023-01-09 Abhinav Kaushal Keshari

Generative AI has been transforming the way we interact with technology and consume content. In the next decade, AI technology will reshape how we create audio content in various media, including music, theater, films, games, podcasts, and…

Sound · Computer Science 2024-11-25 Hao-Wen Dong

Text-to-Image and Text-to-Video AI generation models are revolutionary technologies that use deep learning and natural language processing (NLP) techniques to create images and videos from textual descriptions. This paper investigates…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Aditi Singh

As generative models become ubiquitous, there is a critical need for fine-grained control over the generation process. Yet, while controlled generation methods from prompting to fine-tuning proliferate, a fundamental question remains…

Artificial Intelligence · Computer Science 2026-01-12 Emily Cheng , Carmen Amo Alonso , Federico Danieli , Arno Blaas , Luca Zappella , Pau Rodriguez , Xavier Suau

Adolescence is marked by strong creative impulses but limited strategies for structured expression, often leading to frustration or disengagement. While generative AI lowers technical barriers and delivers efficient outputs, its role in…

Human-Computer Interaction · Computer Science 2025-09-15 Zhejing Hu , Yan Liu , Zhi Zhang , Gong Chen , Bruce X. B. Yu , Junxian Li , Jiannong Cao

Effective music mixing requires technical and creative finesse, but clear communication with the client is crucial. The mixing engineer must grasp the client's expectations, and preferences, and collaborate to achieve the desired sound. The…

Human-Computer Interaction · Computer Science 2023-10-02 Soumya Sai Vanka , Maryam Safi , Jean-Baptiste Rolland , György Fazekas

Text-to-music generation has advanced rapidly, with modern autoregressive and diffusion-based models producing convincing music from natural-language prompts. However, much of this progress relies on large-scale training data and external…

Sound · Computer Science 2026-05-21 Junyoung Koh

Digital advances have transformed the face of automatic music generation since its beginnings at the dawn of computing. Despite the many breakthroughs, issues such as the musical tasks targeted by different machines and the degree to which…

Sound · Computer Science 2018-12-12 Dorien Herremans , Ching-Hua Chuan , Elaine Chew

End-to-end generation of musical audio using deep learning techniques has seen an explosion of activity recently. However, most models concentrate on generating fully mixed music in response to abstract conditioning information. In this…

Natural language processing (NLP) researchers develop models of grammar, meaning and communication based on written text. Due to task and data differences, what is considered text can vary substantially across studies. A conceptual…

Computation and Language · Computer Science 2023-05-18 Ilia Kuznetsov , Iryna Gurevych

Artificial intelligence (AI) is currently based largely on black-box machine learning models which lack interpretability. The field of eXplainable AI (XAI) strives to address this major concern, being critical in high-stakes areas such as…

Artificial Intelligence · Computer Science 2024-06-26 Sean Tull , Robin Lorenz , Stephen Clark , Ilyas Khan , Bob Coecke

This position paper argues that large language models (LLMs) can make cultural context, and therefore human meaning, legible at an unprecedented scale in AI-based sociotechnical systems. We argue that such systems have previously been…

Computation and Language · Computer Science 2026-01-28 Cody Kommers , Drew Hemment , Maria Antoniak , Joel Z. Leibo , Hoyt Long , Emily Robinson , Adam Sobey

Music foundation models possess impressive music generation capabilities. When people compose music, they may infuse their understanding of music into their work, by using notes and intervals to craft melodies, chords to build progressions,…

Sound · Computer Science 2024-10-02 Megan Wei , Michael Freeman , Chris Donahue , Chen Sun

Text-To-Music (TTM) models have recently revolutionized the automatic music generation research field. Specifically, by reaching superior performances to all previous state-of-the-art models and by lowering the technical proficiency needed…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-26 Luca Comanducci , Paolo Bestagini , Stefano Tubaro

Generative models have become adept at producing artifacts such as images, videos, and prose at human-like levels of proficiency. New generative techniques, such as unsupervised neural machine translation (NMT), have recently been applied…

Human-Computer Interaction · Computer Science 2021-04-09 Justin D. Weisz , Michael Muller , Stephanie Houde , John Richards , Steven I. Ross , Fernando Martinez , Mayank Agarwal , Kartik Talamadupula

In this work, we investigate the personalization of text-to-music diffusion models in a few-shot setting. Motivated by recent advances in the computer vision domain, we are the first to explore the combination of pre-trained text-to-audio…