English

TimbreCLIP: Connecting Timbre to Text and Images

Sound 2022-11-22 v1 Machine Learning Audio and Speech Processing

Abstract

We present work in progress on TimbreCLIP, an audio-text cross modal embedding trained on single instrument notes. We evaluate the models with a cross-modal retrieval task on synth patches. Finally, we demonstrate the application of TimbreCLIP on two tasks: text-driven audio equalization and timbre to image generation.

Keywords

Cite

@article{arxiv.2211.11225,
  title  = {TimbreCLIP: Connecting Timbre to Text and Images},
  author = {Nicolas Jonason and Bob L. T. Sturm},
  journal= {arXiv preprint arXiv:2211.11225},
  year   = {2022}
}

Comments

Submitted to AAAI workshop on creative AI across modalities