English

Automatic Music Transcription using Convolutional Neural Networks and Constant-Q transform

Sound 2025-05-08 v1 Artificial Intelligence Machine Learning Multimedia Audio and Speech Processing

Abstract

Automatic music transcription (AMT) is the problem of analyzing an audio recording of a musical piece and detecting notes that are being played. AMT is a challenging problem, particularly when it comes to polyphonic music. The goal of AMT is to produce a score representation of a music piece, by analyzing a sound signal containing multiple notes played simultaneously. In this work, we design a processing pipeline that can transform classical piano audio files in .wav format into a music score representation. The features from the audio signals are extracted using the constant-Q transform, and the resulting coefficients are used as an input to the convolutional neural network (CNN) model.

Keywords

Cite

@article{arxiv.2505.04451,
  title  = {Automatic Music Transcription using Convolutional Neural Networks and Constant-Q transform},
  author = {Yohannis Telila and Tommaso Cucinotta and Davide Bacciu},
  journal= {arXiv preprint arXiv:2505.04451},
  year   = {2025}
}

Comments

6 pages

R2 v1 2026-06-28T23:24:32.405Z