English

Shamela: A Large-Scale Historical Arabic Corpus

Computation and Language 2016-12-30 v1

Abstract

Arabic is a widely-spoken language with a rich and long history spanning more than fourteen centuries. Yet existing Arabic corpora largely focus on the modern period or lack sufficient diachronic information. We develop a large-scale, historical corpus of Arabic of about 1 billion words from diverse periods of time. We clean this corpus, process it with a morphological analyzer, and enhance it by detecting parallel passages and automatically dating undated texts. We demonstrate its utility with selected case-studies in which we show its application to the digital humanities.

Keywords

Cite

@article{arxiv.1612.08989,
  title  = {Shamela: A Large-Scale Historical Arabic Corpus},
  author = {Yonatan Belinkov and Alexander Magidow and Maxim Romanov and Avi Shmidman and Moshe Koppel},
  journal= {arXiv preprint arXiv:1612.08989},
  year   = {2016}
}

Comments

Slightly expanded version of Coling LT4DH workshop paper

R2 v1 2026-06-22T17:36:17.930Z