English

Text-Independent Speaker Verification Using Discrete Audio Tokens

Audio and Speech Processing 2026-07-08 v1

Abstract

Neural audio codecs (NACs) enable efficient audio compression and have achieved success in downstream tasks such as speech synthesis. However, their discrete representations consistently underperform traditional spectral features in automatic speaker verification (ASV). We empirically demonstrate that speaker cues are implicitly preserved in discrete tokens but remain underutilized by conventional ASV training paradigms. To address this, we propose a Cross-Feature Knowledge Distillation (CFKD) framework. By guiding the codec-based student to mimic the embedding space of a strong Fbank-based teacher, CFKD provides structured supervision for effective utilization of speaker information in tokens. Experiments on the VoxCeleb benchmarks show that CFKD substantially improves the ASV performance of codec-based systems, allowing them to approach the accuracy of Fbank-based teacher models and highlighting the potential of discrete audio tokens for diverse speech tasks.

Cite

@article{arxiv.2607.07579,
  title  = {Text-Independent Speaker Verification Using Discrete Audio Tokens},
  author = {Zheng Liang and Junjie Li and Kong Aik Lee},
  journal= {arXiv preprint arXiv:2607.07579},
  year   = {2026}
}

Comments

This paper has been accepted by Interspeech 2026

R2 v1 2026-07-22T20:30:47.621Z