Traditionally, Knowledge Distillation (KD) is used for model compression, often leading to suboptimal performance. In this paper, we evaluate the impact of combining KD loss with alternative pruning techniques, including Low-Rank Factorization (LRF) and l0 regularization, on a conformer-based pre-trained network under the paradigm of Self-Supervised Learning (SSL). We also propose a strategy to jointly prune and train an RNN-T-based ASR model, demonstrating that this approach yields superior performance compared to pruning a pre-trained network first and then using it for ASR training. This approach led to a significant reduction in word error rate: l0 and KD combination achieves the best non-streaming performance, with a 8.9% Relative Word Error Rate (RWER) improvement over the baseline, while LRF and KD combination yields the best results for streaming ASR, improving RWER by 13.4%.
@article{arxiv.2502.05837,
title = {Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models},
author = {Shiva Kumar C and Jitendra Kumar Dhiman and Nagaraj Adiga and Shatrughan Singh},
journal= {arXiv preprint arXiv:2502.05837},
year = {2025}
}
Comments
5 pages, 2 figures, 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing