English
Related papers

Related papers: A Semantic Timbre Dataset for the Electric Guitar

200 papers

Deep learning models define the state-of-the-art in Automatic Drum Transcription (ADT), yet their performance is contingent upon large-scale, paired audio-MIDI datasets, which are scarce. Existing workarounds that use synthetic data often…

Sound · Computer Science 2026-01-15 Pierfrancesco Melucci , Paolo Merialdo , Taketo Akama

Most organisms including humans function by coordinating and integrating sensory signals with motor actions to survive and accomplish desired tasks. Learning these complex sensorimotor mappings proceeds simultaneously and often in an…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-26 Yashish M. Siriwardena , Carol Espy-Wilson , Shihab Shamma

Research on soundscapes has shifted the focus of environmental acoustics from noise levels to the perception of sounds, incorporating contextual factors. Soundscape emotion recognition (SER) models perception using a set of features, with…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-19 Samuel Rey , Luca Martino , Roberto San Millan , Eduardo Morgado

The dominant approach for music representation learning involves the deep unsupervised model family variational autoencoder (VAE). However, most, if not all, viable attempts on this problem have largely been limited to monophonic music.…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-18 Ziyu Wang , Yiyi Zhang , Yixiao Zhang , Junyan Jiang , Ruihan Yang , Junbo Zhao , Gus Xia

The Synesthetic Variational Autoencoder (SynVAE) introduced in this research is able to learn a consistent mapping between visual and auditive sensory modalities in the absence of paired datasets. A quantitative evaluation on MNIST as well…

Computer Vision and Pattern Recognition · Computer Science 2019-09-15 Maximilian Müller-Eberstein , Nanne van Noord

In recent years, Text-to-Audio Generation has achieved remarkable progress, offering sound creators powerful tools to transform textual inspirations into vivid audio. However, existing models predominantly operate directly in the acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-30 Zheqi Dai , Guangyan Zhang , Haolin He , Xiquan Li , Jingyu Li , Chunyat Wu , Yiwen Guo , Qiuqiang Kong

Interpretable machine learning is rapidly becoming a crucial tool for scientific discovery. Among existing approaches, variational autoencoders (VAEs) have shown promise in extracting the hidden physical features of some input data, with no…

High-fidelity text-to-music generation typically relies on massive proprietary datasets and immense computational resources. Existing models often struggle to generate coherent pure musical accompaniments and lack precise, localized…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-19 Huakang Chen , Wenkai Cheng , Guobin Ma , Chunbo Hao , Yuxuan Xia , Mengqi Wei , Zhixian Zhao , Pengcheng Zhu , Hanbing Zhang , Lei Xie

In this work, we address the problem of musical timbre transfer, where the goal is to manipulate the timbre of a sound sample from one instrument to match another instrument while preserving other musical content, such as pitch, rhythm, and…

Sound · Computer Science 2023-10-24 Sicong Huang , Qiyang Li , Cem Anil , Xuchan Bao , Sageev Oore , Roger B. Grosse

A growing number of approaches exist to generate explanations for image classification. However, few of these approaches are subjected to human-subject evaluations, partly because it is challenging to design controlled experiments with…

Artificial Intelligence · Computer Science 2021-05-07 Martin Schuessler , Philipp Weiß , Leon Sixt

This article presents a benchmark study of symbolic piano music classification using the masked language modelling approach of the Bidirectional Encoder Representations from Transformers (BERT). Specifically, we consider two types of MIDI…

Sound · Computer Science 2024-04-16 Yi-Hui Chou , I-Chun Chen , Chin-Jui Chang , Joann Ching , Yi-Hsuan Yang

The fidelity with which neural networks can now generate content such as music presents a scientific opportunity: these systems appear to have learned implicit theories of such content's structure through statistical learning alone. This…

Sound · Computer Science 2026-03-03 Nikhil Singh , Manuel Cherep , Pattie Maes

Existing audio-to-MIDI tools extract notes but discard the timbral characteristics that define an instrument's identity. We present Instrumental, a system that recovers continuous synthesizer parameters from audio by coupling a…

Sound · Computer Science 2026-03-18 Philipp Bogdan

Music information retrieval faces a challenge in modeling contextualized musical concepts formulated by a set of co-occurring tags. In this paper, we investigate the suitability of our recently proposed approach based on a Siamese neural…

Machine Learning · Computer Science 2016-06-08 Ubai Sandouk , Ke Chen

We propose a pre-trained BERT-like model for symbolic music understanding that achieves competitive performance across a wide range of downstream tasks. To achieve this target, we design two novel pre-training objectives, namely token…

Sound · Computer Science 2025-07-08 Jun-You Wang , Li Su

The range of potential applications of acoustic analysis is wide. Classification of sounds, in particular, is a typical machine learning task that received a lot of attention in recent years. The most common approaches to sound…

Sparse Autoencoder (SAE) has emerged as a powerful tool for mechanistic interpretability of large language models. Recent works apply SAE to protein language models (PLMs), aiming to extract and analyze biologically meaningful features from…

Quantitative Methods · Quantitative Biology 2026-01-21 Xiangyu Liu , Haodi Lei , Yi Liu , Yang Liu , Wei Hu

Vocal Percussion Transcription (VPT) is concerned with the automatic detection and classification of vocal percussion sound events, allowing music creators and producers to sketch drum lines on the fly. Classifier algorithms in VPT systems…

Sound · Computer Science 2022-04-12 Alejandro Delgado , Emir Demirel , Vinod Subramanian , Charalampos Saitis , Mark Sandler

In this paper, we generate and control semantically interpretable filters that are directly learned from natural images in an unsupervised fashion. Each semantic filter learns a visually interpretable local structure in conjunction with…

Computer Vision and Pattern Recognition · Computer Science 2019-02-19 Mohit Prabhushankar , Gukyeong Kwon , Dogancan Temel , Ghassan AlRegib

Machine listening systems often rely on fixed taxonomies to organize and label audio data, key for training and evaluating deep neural networks (DNNs) and other supervised algorithms. However, such taxonomies face significant constraints:…

Sound · Computer Science 2024-09-19 Paraskevas Stamatiadis , Michel Olvera , Slim Essid