默认值不够用:德国复数系统的案例研究
cmp-lg
2008-02-03 v1 计算与语言
摘要
德国复数系统已成为语言理论(包括语言学和认知语言学)之间冲突的焦点。我们呈现了使用三种简单分类器的模拟结果——普通最近邻算法、Nosofsky 的“广义语境模型”(GCM)和标准的三层反向传播网络——从德语中名词单数形式的音位表征预测其复数类。虽然这些是极简模型(在架构和输入信息方面),但它们仍做得相当出色。最近邻算法在来自 CELEX 数据库的 24,640 个名词的集合上,以 72% 的准确率预测正确的复数类。对于 8,598 个(非复合)名词的子集,最近邻、GCM 和网络分别在新颖项目上得分 71.0%、75.0% 和 83.5%。此外,它们在该数据集上优于 Marcus 等人(1995)提出的混合模型,即“模式关联器 + 默认规则”模型。
引用
@article{arxiv.cmp-lg/9605020,
title = {Where Defaults Don't Help: the Case of the German Plural System},
author = {Ramin Charles Nakisa and Ulrike Hahn},
journal= {arXiv preprint arXiv:cmp-lg/9605020},
year = {2008}
}
备注
In proceedings of the 18th annual meeting of the Cognitive Science Society. 6 pages, 4 postscript figures. cmp-lg has trouble handling PS fonts under NFSS, so the postscript version is 6 pages in postscript Times Roman and the LaTeX source produces 7 pages of Computer Modern. Just add times to the \usepackage command if your system can handle it