Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis
Abstract
Arabic spans over 30 spoken varieties, yet no open-source text-to-speech system unifies them. Key barriers include substantial cross-dialect lexical and phonological divergence, scarce synthesis-grade data, and the absence of a standardized multi-dialect evaluation benchmark. We present Habibi, a unified-dialectal Arabic TTS framework that addresses all three. Through a multi-step curation pipeline, we repurpose open-source ASR corpora into TTS training data covering 12+ regional dialects. A linguistically-informed curriculum learning strategy - progressing from Modern Standard Arabic to dialectal data - enables robust zero-shot synthesis without text diacritization. We further release the first standardized multi-dialect Arabic TTS benchmark, comprising over 11,000 utterances across 7 dialect subsets with manually verified transcripts. On this benchmark, our unified model matches or surpasses per-dialect specialized models. Both automatic metrics and human evaluations confirm that Habibi is highly competitive with ElevenLabs' Eleven v3 (alpha) in intelligibility, speaker similarity, and naturalness. Extensive ablations (~8,000 H100 GPU hours, 30+ configurations) validate each design choice. We open-source all checkpoints, training and inference code, and benchmark data - the first such release for multi-dialect Arabic TTS - at https://SWivid.github.io/Habibi/ .
Cite
@article{arxiv.2601.13802,
title = {Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis},
author = {Yushen Chen and Junzhe Liu and Yujie Tu and Zhikang Niu and Yuzhe Liang and Chunyu Qiang and Chen Zhang and Kai Yu and Xie Chen},
journal= {arXiv preprint arXiv:2601.13802},
year = {2026}
}