English

DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer ASR

Audio and Speech Processing 2024-03-05 v2 Artificial Intelligence Machine Learning Sound

Abstract

Conformer-based end-to-end models have become ubiquitous these days and are commonly used in both streaming and non-streaming automatic speech recognition (ASR). Techniques like dual-mode and dynamic chunk training helped unify streaming and non-streaming systems. However, there remains a performance gap between streaming with a full and limited past context. To address this issue, we propose the integration of a novel dynamic contextual carry-over mechanism in a state-of-the-art (SOTA) unified ASR system. Our proposed dynamic context Conformer (DCTX-Conformer) utilizes a non-overlapping contextual carry-over mechanism that takes into account both the left context of a chunk and one or more preceding context embeddings. We outperform the SOTA by a relative 25.0% word error rate, with a negligible latency impact due to the additional context embeddings.

Keywords

Cite

@article{arxiv.2306.08175,
  title  = {DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer ASR},
  author = {Goeric Huybrechts and Srikanth Ronanki and Xilai Li and Hadis Nosrati and Sravan Bodapati and Katrin Kirchhoff},
  journal= {arXiv preprint arXiv:2306.08175},
  year   = {2024}
}