中文

一种基于模板的快速自动识别网页正文内容的方法

信息检索 2019-11-27 v1

摘要

搜索引擎已成为浏览互联网信息不可或缺的工具。然而,用户常常被来自不相关网页的冗余结果所困扰。一个原因是搜索引擎也会关注网页中非信息性的块,如广告、导航链接等。在本文中,我们提出一种称为 FastContentExtractor 的快速算法,通过改进 ContentExtractor 算法来自动检测网页中的主要内容块。通过自动识别并存储代表某网站内容块结构的模板,可快速提取该网站新网页的内容块。输出块的层次顺序也得以保持,从而确保提取的内容块与原始顺序一致。

关键词

引用

@article{arxiv.1911.11473,
  title  = {A Fast Template-based Approach to Automatically Identify Primary Text Content of a Web Page},
  author = {Dat Quoc Nguyen and Dai Quoc Nguyen and Son Bao Pham and The Duy Bui},
  journal= {arXiv preprint arXiv:1911.11473},
  year   = {2019}
}

备注

In Proceedings of the 2009 International Conference on Knowledge and Systems Engineering (KSE 2009)