English

Learning Features of Music from Scratch

Machine Learning 2017-04-07 v2 Machine Learning Sound

Abstract

This paper introduces a new large-scale music dataset, MusicNet, to serve as a source of supervision and evaluation of machine learning methods for music research. MusicNet consists of hundreds of freely-licensed classical music recordings by 10 composers, written for 11 instruments, together with instrument/note annotations resulting in over 1 million temporal labels on 34 hours of chamber music performances under various studio and microphone conditions. The paper defines a multi-label classification task to predict notes in musical recordings, along with an evaluation protocol, and benchmarks several machine learning architectures for this task: i) learning from spectrogram features; ii) end-to-end learning with a neural net; iii) end-to-end learning with a convolutional neural net. These experiments show that end-to-end models trained for note prediction learn frequency selective filters as a low-level representation of audio.

Cite

@article{arxiv.1611.09827,
  title  = {Learning Features of Music from Scratch},
  author = {John Thickstun and Zaid Harchaoui and Sham Kakade},
  journal= {arXiv preprint arXiv:1611.09827},
  year   = {2017}
}

Comments

14 pages; camera-ready version; updated experiments and related works; additional MIR metrics (Appendix C)

R2 v1 2026-06-22T17:08:31.382Z