中文

SpEnD:利用搜索引擎发现关联数据 SPARQL 端点

信息检索 2017-04-11 v2

摘要

本研究提出了一种新颖的元爬行(metacrawling)方法,用于发现和监测 Web 上的关联数据(linked data)源。我们在名为 SPARQL Endpoints Discovery (SpEnD) 的原型系统中实现了该方法。SpEnD 首先进行“搜索关键词”发现过程,以寻找关联数据领域特别是 SPARQL 端点的相关关键词。然后,利用这些搜索关键词通过流行的搜索引擎(Google、Bing、Yahoo、Yandex)来查找关联数据源。通过使用该方法,发现了现有端点存储库中当前列出的大部分 SPARQL 端点,以及大量新的 SPARQL 端点。最后,我们开发了一个新的 SPARQL 端点爬虫(SpEC)用于爬行和链接分析。

关键词

引用

@article{arxiv.1608.02761,
  title  = {SpEnD: Linked Data SPARQL Endpoints Discovery Using Search Engines},
  author = {Semih Yumusak and Erdogan Dogdu and Halife Kodaz and Andreas Kamilaris},
  journal= {arXiv preprint arXiv:1608.02761},
  year   = {2017}
}

备注

This paper has been withdrawn by the author due to a crucial illustration error in Figure 2, critical numerical errors in Table 5 and Table 8