English

In-domain SSL pre-training and streaming ASR

Computation and Language 2025-09-16 v1 Artificial Intelligence

Abstract

In this study, we investigate the benefits of domain-specific self-supervised pre-training for both offline and streaming ASR in Air Traffic Control (ATC) environments. We train BEST-RQ models on 4.5k hours of unlabeled ATC data, then fine-tune on a smaller supervised ATC set. To enable real-time processing, we propose using chunked attention and dynamic convolutions, ensuring low-latency inference. We compare these in-domain SSL models against state-of-the-art, general-purpose speech encoders such as w2v-BERT 2.0 and HuBERT. Results show that domain-adapted pre-training substantially improves performance on standard ATC benchmarks, significantly reducing word error rates when compared to models trained on broad speech corpora. Furthermore, the proposed streaming approach further improves word error rate under tighter latency constraints, making it particularly suitable for safety-critical aviation applications. These findings highlight that specializing SSL representations for ATC data is a practical path toward more accurate and efficient ASR systems in real-world operational settings.

Cite

@article{arxiv.2509.12101,
  title  = {In-domain SSL pre-training and streaming ASR},
  author = {Jarod Duret and Salima Mdhaffar and Gaëlle Laperrière and Ryan Whetten and Audrey Galametz and Catherine Kobus and Marion-Cécile Martin and Jo Oleiwan and Yannick Estève},
  journal= {arXiv preprint arXiv:2509.12101},
  year   = {2025}
}

Comments

Accepted to SPECOM 2025

R2 v1 2026-07-01T05:37:13.526Z