English

Explainable Disentanglement on Discrete Speech Representations for Noise-Robust ASR

Computation and Language 2025-10-30 v1

Abstract

Discrete audio representations are gaining traction in speech modeling due to their interpretability and compatibility with large language models, but are not always optimized for noisy or real-world environments. Building on existing works that quantize Whisper embeddings for speech-to-unit modeling, we propose disentangling semantic speech content from background noise in the latent space. Our end-to-end model separates clean speech in the form of codebook tokens, while extracting interpretable noise vectors as quantization residue which are supervised via a lightweight classifier. We show that our approach improves alignment between clean/noisy speech and text, producing speech tokens that display a high degree of noiseinvariance, and improves ASR performance. Keeping Whisper frozen, we show an 82% reduction in error rate compared to Whisper, and 35% improvement over baseline methods on the VBDemand test set. Further analyses show that the learned token space generalizes well to both seen and unseen acoustic conditions.

Keywords

Cite

@article{arxiv.2510.25150,
  title  = {Explainable Disentanglement on Discrete Speech Representations for Noise-Robust ASR},
  author = {Shreyas Gopal and Ashutosh Anshul and Haoyang Li and Yue Heng Yeo and Hexin Liu and Eng Siong Chng},
  journal= {arXiv preprint arXiv:2510.25150},
  year   = {2025}
}

Comments

Awarded Best Student Paper at APSIPA ASC 2025

R2 v1 2026-07-01T07:11:01.211Z