English

Compose & Embellish: Well-Structured Piano Performance Generation via A Two-Stage Approach

Sound 2023-03-08 v4 Artificial Intelligence Multimedia Audio and Speech Processing

Abstract

Even with strong sequence models like Transformers, generating expressive piano performances with long-range musical structures remains challenging. Meanwhile, methods to compose well-structured melodies or lead sheets (melody + chords), i.e., simpler forms of music, gained more success. Observing the above, we devise a two-stage Transformer-based framework that Composes a lead sheet first, and then Embellishes it with accompaniment and expressive touches. Such a factorization also enables pretraining on non-piano data. Our objective and subjective experiments show that Compose & Embellish shrinks the gap in structureness between a current state of the art and real performances by half, and improves other musical aspects such as richness and coherence as well.

Keywords

Cite

@article{arxiv.2209.08212,
  title  = {Compose & Embellish: Well-Structured Piano Performance Generation via A Two-Stage Approach},
  author = {Shih-Lun Wu and Yi-Hsuan Yang},
  journal= {arXiv preprint arXiv:2209.08212},
  year   = {2023}
}

Comments

Accepted to International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2023

R2 v1 2026-06-28T01:29:11.234Z