English

Multi-modal Conditional Bounding Box Regression for Music Score Following

Sound 2021-05-11 v1 Machine Learning Audio and Speech Processing

Abstract

This paper addresses the problem of sheet-image-based on-line audio-to-score alignment also known as score following. Drawing inspiration from object detection, a conditional neural network architecture is proposed that directly predicts x,y coordinates of the matching positions in a complete score sheet image at each point in time for a given musical performance. Experiments are conducted on a synthetic polyphonic piano benchmark dataset and the new method is compared to several existing approaches from the literature for sheet-image-based score following as well as an Optical Music Recognition baseline. The proposed approach achieves new state-of-the-art results and furthermore significantly improves the alignment performance on a set of real-world piano recordings by applying Impulse Responses as a data augmentation technique.

Keywords

Cite

@article{arxiv.2105.04309,
  title  = {Multi-modal Conditional Bounding Box Regression for Music Score Following},
  author = {Florian Henkel and Gerhard Widmer},
  journal= {arXiv preprint arXiv:2105.04309},
  year   = {2021}
}

Comments

Accepted for publication in the Proceedings of the 29th European Signal Processing Conference (EUSIPCO), Dublin, Ireland, 2021

R2 v1 2026-06-24T01:56:33.691Z