English

Scribosermo: Fast Speech-to-Text models for German and other Languages

Computation and Language 2021-10-18 v1 Sound Audio and Speech Processing

Abstract

Recent Speech-to-Text models often require a large amount of hardware resources and are mostly trained in English. This paper presents Speech-to-Text models for German, as well as for Spanish and French with special features: (a) They are small and run in real-time on microcontrollers like a RaspberryPi. (b) Using a pretrained English model, they can be trained on consumer-grade hardware with a relatively small dataset. (c) The models are competitive with other solutions and outperform them in German. In this respect, the models combine advantages of other approaches, which only include a subset of the presented features. Furthermore, the paper provides a new library for handling datasets, which is focused on easy extension with additional datasets and shows an optimized way for transfer-learning new languages using a pretrained model from another language with a similar alphabet.

Keywords

Cite

@article{arxiv.2110.07982,
  title  = {Scribosermo: Fast Speech-to-Text models for German and other Languages},
  author = {Daniel Bermuth and Alexander Poeppel and Wolfgang Reif},
  journal= {arXiv preprint arXiv:2110.07982},
  year   = {2021}
}
R2 v1 2026-06-24T06:54:56.643Z