English

Beat and Downbeat Tracking in Performance MIDI Using an End-to-End Transformer Architecture

Sound 2025-07-02 v1 Computation and Language Multimedia Audio and Speech Processing

Abstract

Beat tracking in musical performance MIDI is a challenging and important task for notation-level music transcription and rhythmical analysis, yet existing methods primarily focus on audio-based approaches. This paper proposes an end-to-end transformer-based model for beat and downbeat tracking in performance MIDI, leveraging an encoder-decoder architecture for sequence-to-sequence translation of MIDI input to beat annotations. Our approach introduces novel data preprocessing techniques, including dynamic augmentation and optimized tokenization strategies, to improve accuracy and generalizability across different datasets. We conduct extensive experiments using the A-MAPS, ASAP, GuitarSet, and Leduc datasets, comparing our model against state-of-the-art hidden Markov models (HMMs) and deep learning-based beat tracking methods. The results demonstrate that our model outperforms existing symbolic music beat tracking approaches, achieving competitive F1-scores across various musical styles and instruments. Our findings highlight the potential of transformer architectures for symbolic beat tracking and suggest future integration with automatic music transcription systems for enhanced music analysis and score generation.

Keywords

Cite

@article{arxiv.2507.00466,
  title  = {Beat and Downbeat Tracking in Performance MIDI Using an End-to-End Transformer Architecture},
  author = {Sebastian Murgul and Michael Heizmann},
  journal= {arXiv preprint arXiv:2507.00466},
  year   = {2025}
}

Comments

Accepted to the 22nd Sound and Music Computing Conference (SMC), 2025