中文
相关论文

相关论文: MIDI-GPT: A Controllable Generative Model for Comp…

200 篇论文

Recent work in the field of symbolic music generation has shown value in using a tokenization based on the GuitarPro format, a symbolic representation supporting guitar expressive attributes, as an input and output representation. We extend…

声音 · 计算机科学 2023-07-12 Jackson Loth , Pedro Sarmento , CJ Carr , Zack Zukowski , Mathieu Barthet

Songs, as a central form of musical art, exemplify the richness of human intelligence and creativity. While recent advances in generative modeling have enabled notable progress in long-form song generation, current systems for full-length…

音频与语音处理 · 电气工程与系统科学 2025-07-25 Huakang Chen , Yuepeng Jiang , Guobin Ma , Chunbo Hao , Shuai Wang , Jixun Yao , Ziqian Ning , Meng Meng , Jian Luan , Lei Xie

Machine generation of symbolic music and digital audio are hot topics but there have been relatively few digital musical instruments that integrate generative AI. Present musical AI tools are not artist centred and do not support…

声音 · 计算机科学 2026-04-28 Charles Patrick Martin

Autoregressive models are now capable of generating high-quality minute-long expressive MIDI piano performances. Even though this progress suggests new tools to assist music composition, we observe that generative algorithms are still not…

声音 · 计算机科学 2021-07-14 Gaëtan Hadjeres , Léopold Crestel

In this paper, we propose SinTra, an auto-regressive sequential generative model that can learn from a single multi-track music segment, to generate coherent, aesthetic, and variable polyphonic music of multi-instruments with an arbitrary…

声音 · 计算机科学 2022-04-22 Qingwei Song , Qiwei Sun , Dongsheng Guo , Haiyong Zheng

We tackle the task of conditional music generation. We introduce MusicGen, a single Language Model (LM) that operates over several streams of compressed discrete music representation, i.e., tokens. Unlike prior work, MusicGen is comprised…

This study proposes a system designed to enumerate the process of collaborative composition among humans, using automatic music composition technology. By integrating multiple Recurrent Neural Network (RNN) models, the system provides an…

声音 · 计算机科学 2024-03-07 So Hirawata , Noriko Otani

We introduce the Latent Fourier Transform (LatentFT), a framework that provides novel frequency-domain controls for generative music models. LatentFT combines a diffusion autoencoder with a latent-space Fourier transform to separate musical…

声音 · 计算机科学 2026-04-21 Mason Wang , Cheng-Zhi Anna Huang

Recent approaches in music generation rely on disentangled representations, often labeled as structure and timbre or local and global, to enable controllable synthesis. Yet the underlying properties of these embeddings remain underexplored.…

We present Composable Diffusion (CoDi), a novel generative model capable of generating any combination of output modalities, such as language, image, video, or audio, from any combination of input modalities. Unlike existing generative AI…

计算机视觉与模式识别 · 计算机科学 2023-05-22 Zineng Tang , Ziyi Yang , Chenguang Zhu , Michael Zeng , Mohit Bansal

Despite deep learning's remarkable advances in style transfer across various domains, generating controllable performance-level musical style transfer for complete symbolically represented musical works remains a challenging area of…

Recently, symbolic music generation with deep learning techniques has witnessed steady improvements. Most works on this topic focus on MIDI representations, but less attention has been paid to symbolic music generation using guitar…

声音 · 计算机科学 2023-02-13 Pedro Sarmento , Adarsh Kumar , Yu-Hua Chen , CJ Carr , Zack Zukowski , Mathieu Barthet

Gesture-driven music generation is an emerging human-computer interaction paradigm for touch-free and expressive musical interaction. However, many existing approaches treat the task as isolated gesture classification or map gestures to…

多媒体 · 计算机科学 2026-04-29 Rathinaraja Jeyaraj , Barathi Subramanian , Kapilya Gangadharan , Anand Paul

This study presents an exploratory evaluation of Music Generation Systems (MGS) within contemporary music production workflows by examining eight open-source systems. The evaluation framework combines technical insights with practical…

音频与语音处理 · 电气工程与系统科学 2025-07-03 Shayan Dadman , Bernt Arild Bremdal , Andreas Bergsland

Representing symbolic music with compound tokens, where each token consists of several different sub-tokens representing a distinct musical feature or attribute, offers the advantage of reducing sequence length. While previous research has…

声音 · 计算机科学 2026-03-17 HaeJun Yoo , Hao-Wen Dong , Jongmin Jung , Dasaem Jeong

This work presents a generative pre-trained transformer (GPT) designed for modeling financial time series. The GPT functions as an order generation engine within a discrete event simulator, enabling realistic replication of limit order book…

交易与市场微观结构 · 定量金融 2024-11-26 Aaron Wheeler , Jeffrey D. Varner

In real-world scenarios, human dialogues are multi-round and diverse. Furthermore, human instructions can be unclear and human responses are unrestricted. Interactive robots face difficulties in understanding human intents and generating…

机器人学 · 计算机科学 2023-08-09 Zhe Zhang , Wei Chai , Jiankun Wang

While recent generative models can produce engaging music, their utility is limited. The variation in the music is often left to chance, resulting in compositions that lack structure. Pieces extending beyond a minute can become incoherent…

声音 · 计算机科学 2023-11-01 Lilac Atassi

Multi-modal music generation, using multiple modalities like text, images, and video alongside musical scores and audio as guidance, is an emerging research area with broad applications. This paper reviews this field, categorizing music…

声音 · 计算机科学 2026-03-09 Shuyu Li , Shulei Ji , Zihao Wang , Songruoyao Wu , Jiaxing Yu , Kejun Zhang

MusicGen is a music generation language model (LM) that can be conditioned on textual descriptions and melodic features. We introduce MusicGen-Chord, which extends this capability by incorporating chord progression features. This model…

声音 · 计算机科学 2024-12-03 Jongmin Jung , Andreas Jansson , Dasaem Jeong