English

CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments

Computation and Language 2025-03-13 v4 Machine Learning Sound Audio and Speech Processing

Abstract

Creating Automatic Speech Recognition (ASR) systems that are robust and resilient to classroom conditions is paramount to the development of AI tools to aid teachers and students. In this work, we study the efficacy of continued pretraining (CPT) in adapting Wav2vec2.0 to the classroom domain. We show that CPT is a powerful tool in that regard and reduces the Word Error Rate (WER) of Wav2vec2.0-based models by upwards of 10%. More specifically, CPT improves the model's robustness to different noises, microphones and classroom conditions.

Keywords

Cite

@article{arxiv.2409.14494,
  title  = {CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments},
  author = {Ahmed Adel Attia and Dorottya Demszky and Tolulope Ogunremi and Jing Liu and Carol Espy-Wilson},
  journal= {arXiv preprint arXiv:2409.14494},
  year   = {2025}
}

Comments

arXiv admin note: substantial text overlap with arXiv:2405.13018

R2 v1 2026-06-28T18:52:57.489Z