面向资源匮乏语言的无监督词义消歧系统
计算与语言
2018-05-01 v1
摘要
在本文中,我们提出 Watasense,一个用于词义消歧的无监督系统。给定句子,该系统根据给定句子与构成目标词词义的同义词集之间的语义相似度,为每个输入词选择最相关的词义。Watasense 有两种运行模式。稀疏模式使用传统的向量空间模型来估计与其上下文对应的最相似词义。而密集模式则使用同义词集嵌入来应对稀疏性问题。我们描述了当前系统的架构,并在三种不同的俄语词汇语义资源上对其进行了评估。根据调整后的兰德指数,我们发现密集模式在所有数据集上都大幅优于稀疏模式。
引用
@article{arxiv.1804.10686,
title = {An Unsupervised Word Sense Disambiguation System for Under-Resourced Languages},
author = {Dmitry Ustalov and Denis Teslenko and Alexander Panchenko and Mikhail Chernoskutov and Chris Biemann and Simone Paolo Ponzetto},
journal= {arXiv preprint arXiv:1804.10686},
year = {2018}
}
备注
In Proceedings of the 11th Conference on Language Resources and Evaluation (LREC 2018). Miyazaki, Japan