English
Related papers

Related papers: Drum Synthesis from Expressive Drum Grids via Neur…

200 papers

Although a variety of transformers have been proposed for symbolic music generation in recent years, there is still little comprehensive study on how specific design choices affect the quality of the generated music. In this work, we…

Deep learning algorithms are increasingly developed for learning to compose music in the form of MIDI files. However, whether such algorithms work well for composing guitar tabs, which are quite different from MIDIs, remain relatively…

Sound · Computer Science 2020-08-05 Yu-Hua Chen , Yu-Hsiang Huang , Wen-Yi Hsiao , Yi-Hsuan Yang

Existing automatic music generation approaches that feature deep learning can be broadly classified into two types: raw audio models and symbolic models. Symbolic models, which train and generate at the note level, are currently the more…

Sound · Computer Science 2018-06-27 Rachel Manzelli , Vijay Thakkar , Ali Siahkamari , Brian Kulis

The advent of neural audio codecs has increased in popularity due to their potential for efficiently modeling audio with transformers. Such advanced codecs represent audio from a highly continuous waveform to low-sampled discrete units. In…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-19 Samir Sadok , Julien Hauret , Éric Bavu

We present the Inverse Drum Machine, a novel approach to Drum Source Separation that leverages an analysis-by-synthesis framework combined with deep learning. Unlike recent supervised methods that require isolated stem recordings for…

Sound · Computer Science 2025-10-01 Bernardo Torres , Geoffroy Peeters , Gael Richard

Recent progress in music generation has been remarkably advanced by the state-of-the-art MusicLM, which comprises a hierarchy of three LMs, respectively, for semantic, coarse acoustic, and fine acoustic modelings. Yet, sampling with the…

Music creation is typically composed of two parts: composing the musical score, and then performing the score with instruments to make sounds. While recent work has made much progress in automatic music generation in the symbolic domain,…

Sound · Computer Science 2018-11-13 Bryan Wang , Yi-Hsuan Yang

Automatic drum transcription (ADT) is traditionally formulated as a discriminative task to predict drum events from audio spectrograms. In this work, we redefine ADT as a conditional generative task and introduce Noise-to-Notes (N2N), a…

Sound · Computer Science 2026-03-06 Michael Yeung , Keisuke Toyama , Toya Teramoto , Shusuke Takahashi , Tamaki Kojima

A method for musical audio synthesis using autoencoding neural networks is proposed. The autoencoder is trained to compress and reconstruct magnitude short-time Fourier transform frames. The autoencoder produces a spectrogram by activating…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-29 Joseph Colonel , Christopher Curro , Sam Keene

We present a system for generating arbitrary, triaxial magnetic waveforms with a spectral content spanning from DC to tens of kHz, a critical capability for quantum control and spin manipulation. To compensate for amplifier-coil dynamics,…

Instrumentation and Detectors · Physics 2026-03-26 Giuseppe Bevilacqua , Valerio Biancalana , Roberto Cecchi

We present a framework based on neural networks to extract music scores directly from polyphonic audio in an end-to-end fashion. Most previous Automatic Music Transcription (AMT) methods seek a piano-roll representation of the pitches, that…

Sound · Computer Science 2019-10-29 Miguel A. Román , Antonio Pertusa , Jorge Calvo-Zaragoza

We consider the problem of learning high-level controls over the global structure of generated sequences, particularly in the context of symbolic music generation with complex language models. In this work, we present the Transformer…

Sound · Computer Science 2020-07-01 Kristy Choi , Curtis Hawthorne , Ian Simon , Monica Dinculescu , Jesse Engel

Emotion-driven melody harmonization aims to generate diverse harmonies for a single melody to convey desired emotions. Previous research found it hard to alter the perceived emotional valence of lead sheets only by harmonizing the same…

Sound · Computer Science 2024-09-26 Jingyue Huang , Yi-Hsuan Yang

Controllable text-to-audio generation aims to synthesize audio from textual descriptions while satisfying user-specified constraints, including event types, temporal sequences, and onset and offset timestamps. This enables precise control…

Sound · Computer Science 2026-02-10 Yisu Liu , Chenxing Li , Wanqian Zhang , Wenfu Wang , Meng Yu , Ruibo Fu , Zheng Lin , Weiping Wang , Dong Yu

Neural audio codecs have made significant strides in efficiently mapping raw audio waveforms into discrete token representations, which are foundational for contemporary audio generative models. However, most existing codecs are optimized…

Photorealistic rendering of dynamic humans is an important ability for telepresence systems, virtual shopping, synthetic data generation, and more. Recently, neural rendering methods, which combine techniques from computer graphics and…

Computer Vision and Pattern Recognition · Computer Science 2020-12-21 Ziyan Wang , Timur Bagautdinov , Stephen Lombardi , Tomas Simon , Jason Saragih , Jessica Hodgins , Michael Zollhöfer

Automatic chord recognition (ACR) extracts time-aligned chord labels from music audio recordings. Despite recent advances, ACR still struggles with oversegmentation, data scarcity, and imbalance, especially in recognizing complex chords…

Sound · Computer Science 2026-04-28 Leekyung Kim , Jonghun Park

While Large Language Models (LLMs) make symbolic music generation increasingly accessible, producing music with distinctive composition and rich expressiveness remains a significant challenge. Many studies have introduced emotion models to…

Sound · Computer Science 2025-11-19 Dengyun Huang , Yonghua Zhu

Neural codecs, comprising an encoder, quantizer, and decoder, enable signal transmission at exceptionally low bitrates. Training these systems requires techniques like the straight-through estimator, soft-to-hard annealing, or statistical…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-10 Wolfgang Mack , Ahmed Mustafa , Rafał Łaganowski , Samer Hijazy

This study explores the extent to which deep learning models can predict groove and its related perceptual dimensions directly from audio signals. We critically examine the effectiveness of seven state-of-the-art deep learning models in…

Sound · Computer Science 2026-03-31 Axel Marmoret , Nicolas Farrugia , Jan Alexander Stupacher