DWUG:四种语言中历时词用法图谱的大型资源
计算与语言
2024-07-09 v3
摘要
词义无论是共时还是历时都难以捕捉。本文描述了基于 100,000 条人类语义邻近判断、在四种不同语言中创建的最大规模分级上下文化历时词义标注资源。我们详尽描述了多轮增量标注过程、用于将用法聚类为义项的聚类算法选择,以及该数据集可能的共时与历时用途。
引用
@article{arxiv.2104.08540,
title = {DWUG: A large Resource of Diachronic Word Usage Graphs in Four Languages},
author = {Dominik Schlechtweg and Nina Tahmasebi and Simon Hengchen and Haim Dubossarsky and Barbara McGillivray},
journal= {arXiv preprint arXiv:2104.08540},
year = {2024}
}
备注
Dominik Schlechtweg, Nina Tahmasebi, Simon Hengchen, Haim Dubossarsky, and Barbara McGillivray. 2021. DWUG: A large Resource of Diachronic Word Usage Graphs in Four Languages. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7079--7091, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics