一种基于模板的快速自动识别网页正文内容的方法
信息检索
2019-11-27 v1
摘要
搜索引擎已成为浏览互联网信息不可或缺的工具。然而,用户常常被来自不相关网页的冗余结果所困扰。一个原因是搜索引擎也会关注网页中非信息性的块,如广告、导航链接等。在本文中,我们提出一种称为 FastContentExtractor 的快速算法,通过改进 ContentExtractor 算法来自动检测网页中的主要内容块。通过自动识别并存储代表某网站内容块结构的模板,可快速提取该网站新网页的内容块。输出块的层次顺序也得以保持,从而确保提取的内容块与原始顺序一致。
引用
@article{arxiv.1911.11473,
title = {A Fast Template-based Approach to Automatically Identify Primary Text Content of a Web Page},
author = {Dat Quoc Nguyen and Dai Quoc Nguyen and Son Bao Pham and The Duy Bui},
journal= {arXiv preprint arXiv:1911.11473},
year = {2019}
}
备注
In Proceedings of the 2009 International Conference on Knowledge and Systems Engineering (KSE 2009)