English

MAJEPPA: Morphing and Assessing in a Unified Piano Performance Space

Audio and Speech Processing 2026-08-11 v1 Multimedia

Abstract

We present MAJEPPA, a self-supervised framework to learn piano performance representations that span the full skill spectrum, from beginner practice sessions to virtuoso concert recordings. We curate the MAJEPPA dataset, comprising ~4,000 annotated recordings across six expertise levels and six recording contexts. We adapt a single pre-trained MIDI autoregressive model with a joint objective: next-token prediction learns score-conditioned performance generation at various skill levels, while InfoNCE and supervised contrastive losses align abstract score and performance representations in a joint embedding space. The proposed model both generates and understands performances in a unified framework. By introducing the EVPMR benchmark, a suite of downstream tasks spanning quality assessment, competition ranking, mistake and technique classification, we evaluate the learnt representations, demonstrating progress towards a real-world model for the piano performance space.

Keywords

Cite

@article{arxiv.2608.11026,
  title  = {MAJEPPA: Morphing and Assessing in a Unified Piano Performance Space},
  author = {Jinwen Zhou and Huan Zhang and Weixi Zhai and Jinhua Liang and Aidan O. T. Hogg and Simon Dixon},
  journal= {arXiv preprint arXiv:2608.11026},
  year   = {2026}
}