English

AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost

Computation and Language 2024-09-20 v1 Sound Audio and Speech Processing

Abstract

We present AutoMode-ASR, a novel framework that effectively integrates multiple ASR systems to enhance the overall transcription quality while optimizing cost. The idea is to train a decision model to select the optimal ASR system for each segment based solely on the audio input before running the systems. We achieve this by ensembling binary classifiers determining the preference between two systems. These classifiers are equipped with various features, such as audio embeddings, quality estimation, and signal properties. Additionally, we demonstrate how using a quality estimator can further improve performance with minimal cost increase. Experimental results show a relative reduction in WER of 16.2%, a cost saving of 65%, and a speed improvement of 75%, compared to using a single-best model for all segments. Our framework is compatible with commercial and open-source black-box ASR systems as it does not require changes in model codes.

Keywords

Cite

@article{arxiv.2409.12476,
  title  = {AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost},
  author = {Ahmet Gündüz and Yunsu Kim and Kamer Ali Yuksel and Mohamed Al-Badrashiny and Thiago Castro Ferreira and Hassan Sawaf},
  journal= {arXiv preprint arXiv:2409.12476},
  year   = {2024}
}

Comments

SPECOM 2024 Conference