English

A Universal Part-of-Speech Tagset

Computation and Language 2015-03-19 v1

Abstract

To facilitate future research in unsupervised induction of syntactic structure and to standardize best-practices, we propose a tagset that consists of twelve universal part-of-speech categories. In addition to the tagset, we develop a mapping from 25 different treebank tagsets to this universal set. As a result, when combined with the original treebank data, this universal tagset and mapping produce a dataset consisting of common parts-of-speech for 22 different languages. We highlight the use of this resource via two experiments, including one that reports competitive accuracies for unsupervised grammar induction without gold standard part-of-speech tags.

Keywords

Cite

@article{arxiv.1104.2086,
  title  = {A Universal Part-of-Speech Tagset},
  author = {Slav Petrov and Dipanjan Das and Ryan McDonald},
  journal= {arXiv preprint arXiv:1104.2086},
  year   = {2015}
}
R2 v1 2026-06-21T17:52:39.237Z