English

Unsupervised Disambiguation of Syncretism in Inflected Lexicons

Computation and Language 2020-02-26 v2

Abstract

Lexical ambiguity makes it difficult to compute various useful statistics of a corpus. A given word form might represent any of several morphological feature bundles. One can, however, use unsupervised learning (as in EM) to fit a model that probabilistically disambiguates word forms. We present such an approach, which employs a neural network to smoothly model a prior distribution over feature bundles (even rare ones). Although this basic model does not consider a token's context, that very property allows it to operate on a simple list of unigram type counts, partitioning each count among different analyses of that unigram. We discuss evaluation metrics for this novel task and report results on 5 languages.

Keywords

Cite

@article{arxiv.1806.03740,
  title  = {Unsupervised Disambiguation of Syncretism in Inflected Lexicons},
  author = {Ryan Cotterell and Christo Kirov and Sabrina J. Mielke and Jason Eisner},
  journal= {arXiv preprint arXiv:1806.03740},
  year   = {2020}
}

Comments

Published at NAACL 2018