English
Related papers

Related papers: STAGE: Stemmed Accompaniment Generation through Pr…

200 papers

The increasing success of audio foundation models across various tasks has led to a growing need for improved interpretability to understand their intricate decision-making processes better. Existing methods primarily focus on explaining…

Sound · Computer Science 2024-10-11 Alican Akman , Qiyang Sun , Björn W. Schuller

We introduce Seed-Music, a suite of music generation systems capable of producing high-quality music with fine-grained style control. Our unified framework leverages both auto-regressive language modeling and diffusion approaches to support…

We present JASCO, a temporally controlled text-to-music generation model utilizing both symbolic and audio-based conditions. JASCO can generate high-quality music samples conditioned on global text descriptions along with fine-grained local…

Sound · Computer Science 2024-06-18 Or Tal , Alon Ziv , Itai Gat , Felix Kreuk , Yossi Adi

In this paper, we explore the tokenized representation of musical scores using the Transformer model to automatically generate musical scores. Thus far, sequence models have yielded fruitful results with note-level (MIDI-equivalent)…

Sound · Computer Science 2021-12-02 Masahiro Suzuki

Building upon Diff-A-Riff, a latent diffusion model for musical instrument accompaniment generation, we present a series of improvements targeting quality, diversity, inference speed, and text-driven control. First, we upgrade the…

Sound · Computer Science 2024-10-31 Javier Nistal , Marco Pasini , Stefan Lattner

Research on automatic music generation has seen great progress due to the development of deep neural networks. However, the generation of multi-instrument music of arbitrary genres still remains a challenge. Existing research either works…

Sound · Computer Science 2018-07-31 Hao-Min Liu , Yi-Hsuan Yang

Generating articulated assets is crucial for robotics, digital twins, and embodied intelligence. Existing generative models often rely on single-view inputs representing closed states, resulting in ambiguous or unrealistic kinematic…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Haowen Wang , Xiaoping Yuan , Fugang Zhang , Rui Jian , Yuanwei Zhu , Xiuquan Qiao , Yakun Huang

Developing text-driven symbolic music generation models remains challenging due to the scarcity of aligned text-music datasets and the unreliability of automated captioning pipelines. While most efforts have focused on MIDI, sheet music…

Performance RNN is a machine-learning system designed primarily for the generation of solo piano performances using an event-based (rather than audio) representation. More specifically, Performance RNN is a long short-term memory (LSTM)…

Sound · Computer Science 2022-02-22 Nicholas Meade , Nicholas Barreyre , Scott C. Lowe , Sageev Oore

A music mashup combines audio elements from two or more songs to create a new work. To reduce the time and effort required to make them, researchers have developed algorithms that predict the compatibility of audio elements. Prior work has…

Sound · Computer Science 2021-03-29 Jiawen Huang , Ju-Chiang Wang , Jordan B. L. Smith , Xuchen Song , Yuxuan Wang

We showcase an unsupervised method that repurposes deep models trained for music generation and music tagging for audio source separation, without any retraining. An audio generation model is conditioned on an input mixture, producing a…

Sound · Computer Science 2021-10-26 Ethan Manilow , Patrick O'Reilly , Prem Seetharaman , Bryan Pardo

Automatically generating natural, diverse and rhythmic human dance movements driven by music is vital for virtual reality and film industries. However, generating dance that naturally follows music remains a challenge, as existing methods…

Multimedia · Computer Science 2025-07-21 Congyi Fan , Jian Guan , Xuanjia Zhao , Dongli Xu , Youtian Lin , Tong Ye , Pengming Feng , Haiwei Pan

Breakthroughs in text-to-music generation models are transforming the creative landscape, equipping musicians with innovative tools for composition and experimentation like never before. However, controlling the generation process to…

Sound · Computer Science 2025-06-19 Teysir Baoueb , Xiaoyu Bie , Xi Wang , Gaël Richard

We propose the Multi-Track Music Machine (MMM), a generative system based on the Transformer architecture that is capable of generating multi-track music. In contrast to previous work, which represents musical material as a single…

Sound · Computer Science 2020-08-24 Jeff Ens , Philippe Pasquier

Music generation has attracted growing interest with the advancement of deep generative models. However, generating music conditioned on textual descriptions, known as text-to-music, remains challenging due to the complexity of musical…

Sound · Computer Science 2025-05-08 Peike Li , Boyu Chen , Yao Yao , Yikai Wang , Allen Wang , Alex Wang

Generating expressive conducting gestures from music is a challenging cross-modal motion synthesis problem: the output must follow long-range musical structure, preserve beat-level synchronization, and remain plausible as a fine-grained 3D…

Sound · Computer Science 2026-05-05 Ke Qiu , Yawen Qin , Tianzhi Jia , Xiaole Yang , Kaimin Wang , Kaixing Yang

We introduce MusicFlow, a cascaded text-to-music generation model based on flow matching. Based on self-supervised representations to bridge between text descriptions and music audios, we construct two flow matching networks to model the…

This study aims to enhance the quality of music generation using Transformers by incorporating meta-information. While Transformer-based approaches are effective at capturing long-term dependencies in musical compositions, the music they…

Sound · Computer Science 2026-05-21 Shinnosuke Taksuka , Hideo Mukai

Music-driven 3D dance generation has attracted increasing attention in recent years, with promising applications in choreography, virtual reality, and creative content creation. Previous research has generated promising realistic dance…

Sound · Computer Science 2026-02-24 Kaixing Yang , Xulong Tang , Ziqiao Peng , Yuxuan Hu , Jun He , Hongyan Liu

Music generation introduces challenging complexities to large language models. Symbolic structures of music often include vertical harmonization as well as horizontal counterpoint, urging various adaptations and enhancements for large-scale…

Sound · Computer Science 2024-07-30 Seungyeon Rhyu , Kichang Yang , Sungjun Cho , Jaehyeon Kim , Kyogu Lee , Moontae Lee