English

Dual-mode ASR: Unify and Improve Streaming ASR with Full-context Modeling

Computation and Language 2021-01-28 v2 Artificial Intelligence Machine Learning Sound Audio and Speech Processing

Abstract

Streaming automatic speech recognition (ASR) aims to emit each hypothesized word as quickly and accurately as possible, while full-context ASR waits for the completion of a full speech utterance before emitting completed hypotheses. In this work, we propose a unified framework, Dual-mode ASR, to train a single end-to-end ASR model with shared weights for both streaming and full-context speech recognition. We show that the latency and accuracy of streaming ASR significantly benefit from weight sharing and joint training of full-context ASR, especially with inplace knowledge distillation during the training. The Dual-mode ASR framework can be applied to recent state-of-the-art convolution-based and transformer-based ASR networks. We present extensive experiments with two state-of-the-art ASR networks, ContextNet and Conformer, on two datasets, a widely used public dataset LibriSpeech and a large-scale dataset MultiDomain. Experiments and ablation studies demonstrate that Dual-mode ASR not only simplifies the workflow of training and deploying streaming and full-context ASR models, but also significantly improves both emission latency and recognition accuracy of streaming ASR. With Dual-mode ASR, we achieve new state-of-the-art streaming ASR results on both LibriSpeech and MultiDomain in terms of accuracy and latency.

Keywords

Cite

@article{arxiv.2010.06030,
  title  = {Dual-mode ASR: Unify and Improve Streaming ASR with Full-context Modeling},
  author = {Jiahui Yu and Wei Han and Anmol Gulati and Chung-Cheng Chiu and Bo Li and Tara N. Sainath and Yonghui Wu and Ruoming Pang},
  journal= {arXiv preprint arXiv:2010.06030},
  year   = {2021}
}

Comments

Accepted in ICLR 2021

R2 v1 2026-06-23T19:17:36.773Z