基于自动获取语言模型的灵活词性标注器
cmp-lg
2008-02-03 v1 计算与语言
摘要
我们提出一种利用统计决策树自动学习上下文约束的算法。随后将获得的约束应用于灵活的词性标注器。该标注器可以使用任意程度的信息:n-gram、自动学习的上下文约束、语言学驱动的手工编写约束等。约束的来源和种类不受限制,语言模型可以方便地扩展,从而提升结果。该标注器已在 WSJ 语料库上进行了测试和评估。
引用
@article{arxiv.cmp-lg/9707003,
title = {A Flexible POS tagger Using an Automatically Acquired Language Model},
author = {Lluis Marquez and Lluis Padro},
journal= {arXiv preprint arXiv:cmp-lg/9707003},
year = {2008}
}
备注
8 pages, aclap.sty, 2 eps figures. Appears in (E)ACL'97