English

DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning

Computation and Language 2024-01-17 v2

Abstract

In this paper, we introduce self-distillation and online clustering for self-supervised speech representation learning (DinoSR) which combines masked language modeling, self-distillation, and online clustering. We show that these concepts complement each other and result in a strong representation learning model for speech. DinoSR first extracts contextualized embeddings from the input audio with a teacher network, then runs an online clustering system on the embeddings to yield a machine-discovered phone inventory, and finally uses the discretized tokens to guide a student network. We show that DinoSR surpasses previous state-of-the-art performance in several downstream tasks, and provide a detailed analysis of the model and the learned discrete units.

Keywords

Cite

@article{arxiv.2305.10005,
  title  = {DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning},
  author = {Alexander H. Liu and Heng-Jui Chang and Michael Auli and Wei-Ning Hsu and James R. Glass},
  journal= {arXiv preprint arXiv:2305.10005},
  year   = {2024}
}
R2 v1 2026-06-28T10:36:46.888Z