English

Sketching the Expression: Flexible Rendering of Expressive Piano Performance with Self-Supervised Learning

Sound 2022-09-07 v2 Multimedia Audio and Speech Processing

Abstract

We propose a system for rendering a symbolic piano performance with flexible musical expression. It is necessary to actively control musical expression for creating a new music performance that conveys various emotions or nuances. However, previous approaches were limited to following the composer's guidelines of musical expression or dealing with only a part of the musical attributes. We aim to disentangle the entire musical expression and structural attribute of piano performance using a conditional VAE framework. It stochastically generates expressive parameters from latent representations and given note structures. In addition, we employ self-supervised approaches that force the latent variables to represent target attributes. Finally, we leverage a two-step encoder and decoder that learn hierarchical dependency to enhance the naturalness of the output. Experimental results show that our system can stably generate performance parameters relevant to the given musical scores, learn disentangled representations, and control musical attributes independently of each other.

Keywords

Cite

@article{arxiv.2208.14867,
  title  = {Sketching the Expression: Flexible Rendering of Expressive Piano Performance with Self-Supervised Learning},
  author = {Seungyeon Rhyu and Sarah Kim and Kyogu Lee},
  journal= {arXiv preprint arXiv:2208.14867},
  year   = {2022}
}

Comments

8 pages, 4 figures, the 23rd International Society for Music Information Retrieval Conference, Bengaluru, India, 2022

R2 v1 2026-06-28T00:29:02.583Z