English

Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization

Computation and Language 2026-07-05 v1 Artificial Intelligence Sound Audio and Speech Processing

Abstract

Unsupervised syllabic tokenization aims to learn discrete syllabic tokens that capture latent linguistic content-related structure from raw speech. Recent syllabic tokenization methods employ teacher-student distillation of the pretrained HuBERT to organize latent speech frame representations into syllabic segments. However, when trained with an utterance-level cross-entropy objective, the model predicts speaker identity rather than linguistic content, thereby compromising the purity of syllabic tokens. To address this problem, we propose a speaker-disentangled syllabic tokenizer that regresses speaker-perturbed student representations toward clean teacher targets within fixed-length chunks. Experimental results demonstrate that our proposed method achieves state-of-the-art performance in syllable boundary detection and syllabic segment clustering. Moreover, a speech language model trained on our syllabic tokens achieves a 7% relative improvement in syntactic and semantic understanding over the phone-level SpiRit-LM.

Cite

@article{arxiv.2607.04064,
  title  = {Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization},
  author = {Ryota Komatsu and Kota Kawakita and Takuma Okamoto and Takahiro Shinozaki},
  journal= {arXiv preprint arXiv:2607.04064},
  year   = {2026}
}

Comments

Accepted by IEEE Open Journal of Signal Processing (OJSP), 10 pages, 4 figures