English

Word frequency-rank relationship in tagged texts

Computation and Language 2021-06-11 v2

Abstract

We analyze the frequency-rank relationship in sub-vocabularies corresponding to three different grammatical classes (nouns, verbs, and others) in a collection of literary works in English, whose words have been automatically tagged according to their grammatical role. Comparing with a null hypothesis which assumes that words belonging to each class are uniformly distributed across the frequency-ranked vocabulary of the whole work, we disclose statistically significant differences between the three classes. This results point to the fact that frequency-rank relationships may reflect linguistic features associated with grammatical function.

Keywords

Cite

@article{arxiv.2102.10992,
  title  = {Word frequency-rank relationship in tagged texts},
  author = {A. Chacoma and D. H. Zanette},
  journal= {arXiv preprint arXiv:2102.10992},
  year   = {2021}
}
R2 v1 2026-06-23T23:23:56.844Z