English

Towards Matching Phones and Speech Representations

Computation and Language 2023-10-27 v1 Machine Learning Sound Audio and Speech Processing

Abstract

Learning phone types from phone instances has been a long-standing problem, while still being open. In this work, we revisit this problem in the context of self-supervised learning, and pose it as the problem of matching cluster centroids to phone embeddings. We study two key properties that enable matching, namely, whether cluster centroids of self-supervised representations reduce the variability of phone instances and respect the relationship among phones. We then use the matching result to produce pseudo-labels and introduce a new loss function for improving self-supervised representations. Our experiments show that the matching result captures the relationship among phones. Training the new loss function jointly with the regular self-supervised losses, such as APC and CPC, significantly improves the downstream phone classification.

Keywords

Cite

@article{arxiv.2310.17558,
  title  = {Towards Matching Phones and Speech Representations},
  author = {Gene-Ping Yang and Hao Tang},
  journal= {arXiv preprint arXiv:2310.17558},
  year   = {2023}
}

Comments

Accepted to ASRU 2023

R2 v1 2026-06-28T13:02:59.845Z