English

Unsupervised Improvement of Audio-Text Cross-Modal Representations

Sound 2023-08-02 v3 Machine Learning Audio and Speech Processing

Abstract

Recent advances in using language models to obtain cross-modal audio-text representations have overcome the limitations of conventional training approaches that use predefined labels. This has allowed the community to make progress in tasks like zero-shot classification, which would otherwise not be possible. However, learning such representations requires a large amount of human-annotated audio-text pairs. In this paper, we study unsupervised approaches to improve the learning framework of such representations with unpaired text and audio. We explore domain-unspecific and domain-specific curation methods to create audio-text pairs that we use to further improve the model. We also show that when domain-specific curation is used in conjunction with a soft-labeled contrastive loss, we are able to obtain significant improvement in terms of zero-shot classification performance on downstream sound event classification or acoustic scene classification tasks.

Keywords

Cite

@article{arxiv.2305.01864,
  title  = {Unsupervised Improvement of Audio-Text Cross-Modal Representations},
  author = {Zhepei Wang and Cem Subakan and Krishna Subramani and Junkai Wu and Tiago Tavares and Fabio Ayres and Paris Smaragdis},
  journal= {arXiv preprint arXiv:2305.01864},
  year   = {2023}
}

Comments

Accepted to WASPAA 2023

R2 v1 2026-06-28T10:24:07.223Z