English

SV-Mixer: Replacing the Transformer Encoder with Lightweight MLPs for Self-Supervised Model Compression in Speaker Verification

Audio and Speech Processing 2025-09-18 v1

Abstract

Self-supervised learning (SSL) has pushed speaker verification accuracy close to state-of-the-art levels, but the Transformer backbones used in most SSL encoders hinder on-device and real-time deployment. Prior compression work trims layer depth or width yet still inherits the quadratic cost of self-attention. We propose SV-Mixer, the first fully MLP-based student encoder for SSL distillation. SV-Mixer replaces Transformer with three lightweight modules: Multi-Scale Mixing for multi-resolution temporal features, Local-Global Mixing for frame-to-utterance context, and Group Channel Mixing for spectral subspaces. Distilled from WavLM, SV-Mixer outperforms a Transformer student by 14.6% while cutting parameters and GMACs by over half, and at 75% compression, it closely matches the teacher's performance. Our results show that attention-free SSL students can deliver teacher-level accuracy with hardware-friendly footprints, opening the door to robust on-device speaker verification.

Cite

@article{arxiv.2509.14136,
  title  = {SV-Mixer: Replacing the Transformer Encoder with Lightweight MLPs for Self-Supervised Model Compression in Speaker Verification},
  author = {Jungwoo Heo and Hyun-seo Shin and Chan-yeong Lim and Kyo-won Koo and Seung-bin Kim and Jisoo Son and Ha-Jin Yu},
  journal= {arXiv preprint arXiv:2509.14136},
  year   = {2025}
}

Comments

8 pages, 5 figures, accepted at IEEE ASRU 2025

R2 v1 2026-07-01T05:42:14.494Z