English

Semi-Automatic Data Annotation, POS Tagging and Mildly Context-Sensitive Disambiguation: the eXtended Revised AraMorph (XRAM)

Computation and Language 2016-03-08 v1 Information Retrieval

Abstract

An extended, revised form of Tim Buckwalter's Arabic lexical and morphological resource AraMorph, eXtended Revised AraMorph (henceforth XRAM), is presented which addresses a number of weaknesses and inconsistencies of the original model by allowing a wider coverage of real-world Classical and contemporary (both formal and informal) Arabic texts. Building upon previous research, XRAM enhancements include (i) flag-selectable usage markers, (ii) probabilistic mildly context-sensitive POS tagging, filtering, disambiguation and ranking of alternative morphological analyses, (iii) semi-automatic increment of lexical coverage through extraction of lexical and morphological information from existing lexical resources. Testing of XRAM through a front-end Python module showed a remarkable success level.

Keywords

Cite

@article{arxiv.1603.01833,
  title  = {Semi-Automatic Data Annotation, POS Tagging and Mildly Context-Sensitive Disambiguation: the eXtended Revised AraMorph (XRAM)},
  author = {Giuliano Lancioni and Valeria Pettinari and Laura Garofalo and Marta Campanelli and Ivana Pepe and Simona Olivieri and Ilaria Cicola},
  journal= {arXiv preprint arXiv:1603.01833},
  year   = {2016}
}
R2 v1 2026-06-22T13:04:42.147Z