English

ATST: Audio Representation Learning with Teacher-Student Transformer

Audio and Speech Processing 2023-06-08 v3 Artificial Intelligence Sound

Abstract

Self-supervised learning (SSL) learns knowledge from a large amount of unlabeled data, and then transfers the knowledge to a specific problem with a limited number of labeled data. SSL has achieved promising results in various domains. This work addresses the problem of segment-level general audio SSL, and proposes a new transformer-based teacher-student SSL model, named ATST. A transformer encoder is developed on a recently emerged teacher-student baseline scheme, which largely improves the modeling capability of pre-training. In addition, a new strategy for positive pair creation is designed to fully leverage the capability of transformer. Extensive experiments have been conducted, and the proposed model achieves the new state-of-the-art results on almost all of the downstream tasks.

Keywords

Cite

@article{arxiv.2204.12076,
  title  = {ATST: Audio Representation Learning with Teacher-Student Transformer},
  author = {Xian Li and Xiaofei Li},
  journal= {arXiv preprint arXiv:2204.12076},
  year   = {2023}
}

Comments

INTERSPEECH2022(Accepted)

R2 v1 2026-06-24T10:58:35.241Z