Constrained CTC Decoding for Efficient Diacritic Restoration
Abstract
In this work, we address diacritic restoration for Arabic speech transcripts. Most speech data are undiacritized, limiting the ability of modeling fine-grained phonological distinctions. The speech modality has recently been explored as a way to complement text-based diacritic restoration efforts. We propose an efficient non-autoregressive approach for speech-to-text diacritization based on Connectionist Temporal Classification (CTC). Our method incorporates hard constraints during decoding by constructing a character-level diacritization lattice from an undiacritized transcript and restricting hypotheses to valid diacritized realizations. We evaluate on Classical Arabic and Modern Standard Arabic test sets (namely, ArVoice and ClArTTS) against a more computationally-complex multi-modal diacritic restoration baseline, and show statistically significant reductions in diacritic error rates in both, demonstrating that the proposed approach offers both performance and efficiency gains.
Cite
@article{arxiv.2607.18946,
title = {Constrained CTC Decoding for Efficient Diacritic Restoration},
author = {Rufael Marew and Amr Keleg and Hanan Aldarmaki},
journal= {arXiv preprint arXiv:2607.18946},
year = {2026}
}
Comments
Accepted at Interspeech 2026