English

AraS2P: Arabic Speech-to-Phonemes System

Computation and Language 2025-09-30 v1

Abstract

This paper describes AraS2P, our speech-to-phonemes system submitted to the Iqra'Eval 2025 Shared Task. We adapted Wav2Vec2-BERT via Two-Stage training strategy. In the first stage, task-adaptive continue pretraining was performed on large-scale Arabic speech-phonemes datasets, which were generated by converting the Arabic text using the MSA Phonetiser. In the second stage, the model was fine-tuned on the official shared task data, with additional augmentation from XTTS-v2-synthesized recitations featuring varied Ayat segments, speaker embeddings, and textual perturbations to simulate possible human errors. The system ranked first on the official leaderboard, demonstrating that phoneme-aware pretraining combined with targeted augmentation yields strong performance in phoneme-level mispronunciation detection.

Keywords

Cite

@article{arxiv.2509.23504,
  title  = {AraS2P: Arabic Speech-to-Phonemes System},
  author = {Bassam Matar and Mohamed Fayed and Ayman Khalafallah},
  journal= {arXiv preprint arXiv:2509.23504},
  year   = {2025}
}
R2 v1 2026-07-01T06:01:34.344Z