中文
相关论文

相关论文: The Interpretation Gap in Text-to-Music Generation…

200 篇论文

Generative AI has demonstrated impressive performance in various fields, among which speech synthesis is an interesting direction. With the diffusion model as the most popular generative model, numerous works have attempted two active…

In this paper, we consider a novel research problem: music-to-text synaesthesia. Different from the classical music tagging problem that classifies a music recording into pre-defined categories, music-to-text synaesthesia aims to generate…

音频与语音处理 · 电气工程与系统科学 2023-05-09 Zhihuan Kuang , Shi Zong , Jianbing Zhang , Jiajun Chen , Hongfu Liu

Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for automatically generating music videos from any…

This paper is a survey and an analysis of different ways of using deep learning (deep artificial neural networks) to generate musical content. We propose a methodology based on five dimensions for our analysis: Objective - What musical…

声音 · 计算机科学 2019-08-09 Jean-Pierre Briot , Gaëtan Hadjeres , François-David Pachet

In the task of generating music, the art factor plays a big role and is a great challenge for AI. Previous work involving adversarial training to produce new music pieces and modeling the compatibility of variety in music (beats, tempo,…

声音 · 计算机科学 2023-01-09 Abhinav Kaushal Keshari

Generative AI has been transforming the way we interact with technology and consume content. In the next decade, AI technology will reshape how we create audio content in various media, including music, theater, films, games, podcasts, and…

声音 · 计算机科学 2024-11-25 Hao-Wen Dong

Text-to-Image and Text-to-Video AI generation models are revolutionary technologies that use deep learning and natural language processing (NLP) techniques to create images and videos from textual descriptions. This paper investigates…

计算机视觉与模式识别 · 计算机科学 2023-11-14 Aditi Singh

As generative models become ubiquitous, there is a critical need for fine-grained control over the generation process. Yet, while controlled generation methods from prompting to fine-tuning proliferate, a fundamental question remains…

人工智能 · 计算机科学 2026-01-12 Emily Cheng , Carmen Amo Alonso , Federico Danieli , Arno Blaas , Luca Zappella , Pau Rodriguez , Xavier Suau

Adolescence is marked by strong creative impulses but limited strategies for structured expression, often leading to frustration or disengagement. While generative AI lowers technical barriers and delivers efficient outputs, its role in…

人机交互 · 计算机科学 2025-09-15 Zhejing Hu , Yan Liu , Zhi Zhang , Gong Chen , Bruce X. B. Yu , Junxian Li , Jiannong Cao

Effective music mixing requires technical and creative finesse, but clear communication with the client is crucial. The mixing engineer must grasp the client's expectations, and preferences, and collaborate to achieve the desired sound. The…

人机交互 · 计算机科学 2023-10-02 Soumya Sai Vanka , Maryam Safi , Jean-Baptiste Rolland , György Fazekas

Text-to-music generation has advanced rapidly, with modern autoregressive and diffusion-based models producing convincing music from natural-language prompts. However, much of this progress relies on large-scale training data and external…

声音 · 计算机科学 2026-05-21 Junyoung Koh

Digital advances have transformed the face of automatic music generation since its beginnings at the dawn of computing. Despite the many breakthroughs, issues such as the musical tasks targeted by different machines and the degree to which…

声音 · 计算机科学 2018-12-12 Dorien Herremans , Ching-Hua Chuan , Elaine Chew

End-to-end generation of musical audio using deep learning techniques has seen an explosion of activity recently. However, most models concentrate on generating fully mixed music in response to abstract conditioning information. In this…

Natural language processing (NLP) researchers develop models of grammar, meaning and communication based on written text. Due to task and data differences, what is considered text can vary substantially across studies. A conceptual…

计算与语言 · 计算机科学 2023-05-18 Ilia Kuznetsov , Iryna Gurevych

Artificial intelligence (AI) is currently based largely on black-box machine learning models which lack interpretability. The field of eXplainable AI (XAI) strives to address this major concern, being critical in high-stakes areas such as…

人工智能 · 计算机科学 2024-06-26 Sean Tull , Robin Lorenz , Stephen Clark , Ilyas Khan , Bob Coecke

This position paper argues that large language models (LLMs) can make cultural context, and therefore human meaning, legible at an unprecedented scale in AI-based sociotechnical systems. We argue that such systems have previously been…

计算与语言 · 计算机科学 2026-01-28 Cody Kommers , Drew Hemment , Maria Antoniak , Joel Z. Leibo , Hoyt Long , Emily Robinson , Adam Sobey

Music foundation models possess impressive music generation capabilities. When people compose music, they may infuse their understanding of music into their work, by using notes and intervals to craft melodies, chords to build progressions,…

声音 · 计算机科学 2024-10-02 Megan Wei , Michael Freeman , Chris Donahue , Chen Sun

Text-To-Music (TTM) models have recently revolutionized the automatic music generation research field. Specifically, by reaching superior performances to all previous state-of-the-art models and by lowering the technical proficiency needed…

音频与语音处理 · 电气工程与系统科学 2024-09-26 Luca Comanducci , Paolo Bestagini , Stefano Tubaro

Generative models have become adept at producing artifacts such as images, videos, and prose at human-like levels of proficiency. New generative techniques, such as unsupervised neural machine translation (NMT), have recently been applied…

In this work, we investigate the personalization of text-to-music diffusion models in a few-shot setting. Motivated by recent advances in the computer vision domain, we are the first to explore the combination of pre-trained text-to-audio…