English

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation

Audio and Speech Processing 2025-05-27 v1

Abstract

Recent advancements in speech encoders have drawn attention due to their integration with Large Language Models for various speech tasks. While most research has focused on either causal or full-context speech encoders, there's limited exploration to effectively handle both streaming and non-streaming applications, while achieving state-of-the-art performance. We introduce DuRep, a Dual-mode Speech Representation learning setup, which enables a single speech encoder to function efficiently in both offline and online modes without additional parameters or mode-specific adjustments, across downstream tasks. DuRep-200M, our 200M parameter dual-mode encoder, achieves 12% and 11.6% improvements in streaming and non-streaming modes, over baseline encoders on Multilingual ASR. Scaling this approach to 2B parameters, DuRep-2B sets new performance benchmarks across ASR and non-ASR tasks. Our analysis reveals interesting trade-offs between acoustic and semantic information across encoder layers.

Keywords

Cite

@article{arxiv.2505.19774,
  title  = {DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation},
  author = {Prabash Reddy Male and Swayambhu Nath Ray and Harish Arsikere and Akshat Jaiswal and Prakhar Swarup and Prantik Sen and Debmalya Chakrabarty and K V Vijay Girish and Nikhil Bhave and Frederick Weber and Sambuddha Bhattacharya and Sri Garimella},
  journal= {arXiv preprint arXiv:2505.19774},
  year   = {2025}
}
R2 v1 2026-07-01T02:39:01.449Z