English

wav2letter++: The Fastest Open-source Speech Recognition System

Computation and Language 2020-02-25 v1

Abstract

This paper introduces wav2letter++, the fastest open-source deep learning speech recognition framework. wav2letter++ is written entirely in C++, and uses the ArrayFire tensor library for maximum efficiency. Here we explain the architecture and design of the wav2letter++ system and compare it to other major open-source speech recognition systems. In some cases wav2letter++ is more than 2x faster than other optimized frameworks for training end-to-end neural networks for speech recognition. We also show that wav2letter++'s training times scale linearly to 64 GPUs, the highest we tested, for models with 100 million parameters. High-performance frameworks enable fast iteration, which is often a crucial factor in successful research and model tuning on new datasets and tasks.

Keywords

Cite

@article{arxiv.1812.07625,
  title  = {wav2letter++: The Fastest Open-source Speech Recognition System},
  author = {Vineel Pratap and Awni Hannun and Qiantong Xu and Jeff Cai and Jacob Kahn and Gabriel Synnaeve and Vitaliy Liptchinsky and Ronan Collobert},
  journal= {arXiv preprint arXiv:1812.07625},
  year   = {2020}
}
R2 v1 2026-06-23T06:46:57.323Z