中文

利用自然语言处理构建塞尔维亚语明喻当代语料库

计算与语言 2018-11-27 v1 人工智能 计算机与社会 机器学习

摘要

明喻是一种通过使用连接词比较两事物的修辞格,但其比较并非按字面理解。它们常用于日常交流,也是语言文化遗产的一部分。本文提出一种利用文本挖掘与机器学习技术从万维网半自动收集明喻的方法。我们通过从互联网收集 442 个明喻并加入由 Vuk Stefanovic Karadzic 收集的含 333 个明喻的现有语料库,扩展了该语料库。我们还引入了众包来收集修辞格,这有助于我们构建包含 787 个唯一明喻的语料库。

关键词

引用

@article{arxiv.1811.10422,
  title  = {Creating a contemporary corpus of similes in Serbian by using natural language processing},
  author = {Nikola Milosevic and Goran Nenadic},
  journal= {arXiv preprint arXiv:1811.10422},
  year   = {2018}
}

备注

15 pages, submitted to journal Slovo, however, later withdrawn to correct. Additional work was not done on it, so it is still waiting to be extended. Output of the system can be seen here: http://ezbirka.starisloveni.com/. arXiv admin note: text overlap with arXiv:1605.06319