Ossetic-COT: Designing a morphologically annotated corpus and morphological analyzer for Ossetic
Computation and Language
2026-07-06 v1
Abstract
In this work we present the first morphologically annotated corpus for Iron Ossetic that conforms to the Universal Dependencies schema. The corpus includes 5454 manually annotated sentences from the Iron Ossetic Corpus of Oral Texts, containing 74032 tokens. We use this corpus to train a BERT-based morphological analyzer. The analyzer achieves tag accuracy of 95.60%.
Cite
@article{arxiv.2607.04895,
title = {Ossetic-COT: Designing a morphologically annotated corpus and morphological analyzer for Ossetic},
author = {Anna Shatskikh and Alexey Sorokin},
journal= {arXiv preprint arXiv:2607.04895},
year = {2026}
}
Comments
12 pages