中文
相关论文

相关论文: Break-the-Beat! Controllable MIDI-to-Drum Audio Sy…

200 篇论文

Modern digital music production typically involves combining numerous acoustic elements to compile a piece of music. Important types of such elements are drum samples, which determine the characteristics of the percussive components of the…

声音 · 计算机科学 2022-08-03 Stefan Lattner

Compared with ample visual-text pre-training research, few works explore audio-text pre-training, mostly due to the lack of sufficient parallel audio-text data. Most existing methods incorporate the visual modality as a pivot for audio-text…

声音 · 计算机科学 2024-03-06 Xuenan Xu , Zhiling Zhang , Zelin Zhou , Pingyue Zhang , Zeyu Xie , Mengyue Wu , Kenny Q. Zhu

Text-to-music models have revolutionized the creative landscape, offering new possibilities for music creation. Yet their integration into musicians workflows remains underexplored. This paper presents a case study on how TTM models impact…

音频与语音处理 · 电气工程与系统科学 2025-09-30 Francesca Ronchini , Luca Comanducci , Simone Marcucci , Fabio Antonacci

Drum one-shot samples are crucial for music production, particularly in sound design and electronic music. This paper introduces Drum One-Shot Extraction, a task in which the goal is to extract drum one-shots that are present in the music…

音频与语音处理 · 电气工程与系统科学 2025-04-28 Suntae Hwang , Seonghyeon Kang , Kyungsu Kim , Semin Ahn , Kyogu Lee

Presented is a method of generating a full drum kit part for a provided kick-drum sequence. A sequence to sequence neural network model used in natural language translation was adopted to encode multiple musical styles and an online survey…

声音 · 计算机科学 2017-06-30 P. Hutchings

Loops, seamlessly repeatable musical segments, are a cornerstone of modern music production. Contemporary artists often mix and match various sampled or pre-recorded loops based on musical criteria such as rhythm, harmony and timbral…

声音 · 计算机科学 2021-05-24 Pritish Chandna , António Ramires , Xavier Serra , Emilia Gómez

Most recent advances in audio dereverberation focus almost exclusively on speech, leaving percussive and drum signals largely unexplored despite their importance in music production. Percussive dereverberation poses distinct challenges due…

声音 · 计算机科学 2026-05-12 Dimos Makris , András Barják , Maximos Kaliakatsos-Papakostas

Dance, as an art form, fundamentally hinges on the precise synchronization with musical beats. However, achieving aesthetically pleasing dance sequences from music is challenging, with existing methods often falling short in controllability…

图形学 · 计算机科学 2024-07-11 Zikai Huang , Xuemiao Xu , Cheng Xu , Huaidong Zhang , Chenxi Zheng , Jing Qin , Shengfeng He

Generative artificial intelligence models can be a valuable aid to music composition and live performance, both to aid the professional musician and to help democratize the music creation process for hobbyists. Here we present a novel…

声音 · 计算机科学 2022-09-22 Ignacio J. Tripodi

Voice conversion aims to transform source speech into a different target voice. However, typical voice conversion systems do not account for rhythm, which is an important factor in the perception of speaker identity. To bridge this gap, we…

音频与语音处理 · 电气工程与系统科学 2023-07-13 Benjamin van Niekerk , Marc-André Carbonneau , Herman Kamper

DeepDrummer is a drum loop generation tool that uses active learning to learn the preferences (or current artistic intentions) of a human user from a small number of interactions. The principal goal of this tool is to enable an efficient…

机器学习 · 计算机科学 2020-08-28 Guillaume Alain , Maxime Chevalier-Boisvert , Frederic Osterrath , Remi Piche-Taillefer

Text-to-music (TTM) generation, which converts textual descriptions into audio, opens up innovative avenues for multimedia creation. Achieving high quality and diversity in this process demands extensive, high-quality data, which are often…

声音 · 计算机科学 2025-06-18 Chang Li , Ruoyu Wang , Lijuan Liu , Jun Du , Yixuan Sun , Zilu Guo , Zhenrong Zhang , Yuan Jiang , Jianqing Gao , Feng Ma

Recent advances in Diffusion Transformers (DiTs) have enabled high-quality joint audio-video generation, producing videos with synchronized audio within a single model. However, existing controllable generation frameworks are typically…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Liyang Li , Wen Wang , Canyu Zhao , Tianjian Feng , Zhiyue Zhao , Hao Chen , Chunhua Shen

Automatic Music Transcription (AMT) is a vital technology in the field of music information processing. Despite recent enhancements in performance due to machine learning techniques, current methods typically attain high accuracy in domains…

声音 · 计算机科学 2024-07-04 Gakusei Sato , Taketo Akama

A crucial aspect for the successful deployment of audio-based models "in-the-wild" is the robustness to the transformations introduced by heterogeneous acquisition conditions. In this work, we propose a method to perform one-shot microphone…

声音 · 计算机科学 2020-10-20 Zalán Borsos , Yunpeng Li , Beat Gfeller , Marco Tagliasacchi

Recent advancements in neural audio codecs have enabled the use of tokenized audio representations in various audio generation tasks, such as text-to-speech, text-to-audio, and text-to-music generation. Leveraging this approach, we propose…

声音 · 计算机科学 2025-02-14 Kyungsu Kim , Junghyun Koo , Sungho Lee , Haesun Joung , Kyogu Lee

The intersection between poetry and music provides an interesting case for computational creativity, yet remains relatively unexplored. This paper explores the integration of poetry and music through the lens of beat patterns, investigating…

计算与语言 · 计算机科学 2025-02-06 Mohamad Elzohbi , Richard Zhao

This paper presents a novel approach to neural instrument sound synthesis using a two-stage semi-supervised learning framework capable of generating pitch-accurate, high-quality music samples from an expressive timbre latent space. Existing…

声音 · 计算机科学 2025-10-07 Christian Limberg , Fares Schulz , Zhe Zhang , Stefan Weinzierl

In this work, we investigate the personalization of text-to-music diffusion models in a few-shot setting. Motivated by recent advances in the computer vision domain, we are the first to explore the combination of pre-trained text-to-audio…

Recent advancements in music source separation have significantly progressed, particularly in isolating vocals, drums, and bass elements from mixed tracks. These developments owe much to the creation and use of large-scale, multitrack…

音频与语音处理 · 电气工程与系统科学 2025-02-18 Jaime Garcia-Martinez , David Diaz-Guerra , Archontis Politis , Tuomas Virtanen , Julio J. Carabias-Orti , Pedro Vera-Candeas