中文
相关论文

相关论文: Extraction of Relevant Images for Boilerplate Remo…

200 篇论文

The extraction of main content from web pages is an important task for numerous applications, ranging from usability aspects, like reader views for news articles in web browsers, to information retrieval or natural language processing.…

机器学习 · 计算机科学 2020-04-30 Jurek Leonhardt , Avishek Anand , Megha Khosla

Boilerplate removal refers to the problem of removing noisy content from a webpage such as ads and extracting relevant content that can be used by various services. This can be useful in several features in web browsers such as ad blocking,…

机器学习 · 统计学 2019-11-11 Joy Bose , Sumanta Mukherjee

Web pages are a valuable source of information for many natural language processing and information retrieval tasks. Extracting the main content from those documents is essential for the performance of derived applications. To address this…

信息检索 · 计算机科学 2018-03-28 Thijs Vogels , Octavian-Eugen Ganea , Carsten Eickhoff

Extracting main content from web pages provides primary informative blocks that remove a web page's minor areas like navigation menu, ads, and site templates. The main content extraction has various applications: information retrieval,…

信息检索 · 计算机科学 2022-01-26 Geunseong Jung , Sungjae Han , Hansung Kim , Kwanguk Kim , Jaehyuk Cha

Template detection and content extraction are two of the main areas of information retrieval applied to the Web. They perform different analyses over the structure and content of webpages to extract some part of the document. However, their…

信息检索 · 计算机科学 2022-07-19 Julián Alarte , Josep Silva

Search engines have become an indispensable tool for browsing information on the Internet. The user, however, is often annoyed by redundant results from irrelevant Web pages. One reason is because search engines also look at non-informative…

信息检索 · 计算机科学 2019-11-27 Dat Quoc Nguyen , Dai Quoc Nguyen , Son Bao Pham , The Duy Bui

We present a hierarchical neural network model called SemText to detect HTML boilerplate based on a novel semantic representation of HTML tags, class names, and text blocks. We train SemText on three published datasets of news webpages and…

计算与语言 · 计算机科学 2022-03-10 Hao Zhang , Jie Wang

Web templates are one of the main development resources for website engineers. Templates allow them to increase productivity by plugin content into already formatted and prepared pagelets. For the final user templates are also useful,…

信息检索 · 计算机科学 2015-01-12 Julián Alarte , David Insa , Josep Silva , Salvador Tamarit

Labeling objects at a subordinate level typically requires expert knowledge, which is not always available when using random annotators. As such, learning directly from web images for fine-grained recognition has attracted broad attention.…

计算机视觉与模式识别 · 计算机科学 2021-01-26 Huafeng Liu , Chuanyi Zhang , Yazhou Yao , Xiushen Wei , Fumin Shen , Jian Zhang , Zhenmin Tang

Searching is an important tool of information gathering, if information is in the form of picture than it play a major role to take quick action and easy to memorize. This is a human tendency to retain more picture than text. The complexity…

信息检索 · 计算机科学 2011-12-12 Anamika Sharma

Web images come in hand with valuable contextual information. Although this information has long been mined for various uses such as image annotation, clustering of images, inference of image semantic content, etc., insufficient attention…

多媒体 · 计算机科学 2020-05-21 F. Fauzi , H. J. Long , M. Belkhatir

Accuracy is one of the basic principles of journalism. However, it is increasingly hard to manage due to the diversity of news media. Some editors of online news tend to use catchy headlines which trick readers into clicking. These…

计算与语言 · 计算机科学 2017-08-30 Wei Wei , Xiaojun Wan

Template extraction is the process of isolating the template of a given webpage. It is widely used in several disciplines, including webpages development, content extraction, block detection, and webpages indexing. One of the main goals of…

信息检索 · 计算机科学 2014-09-10 Julián Alarte , David Insa , Josep Silva , Salvador Tamarit

The main information of a webpage is usually mixed between menus, advertisements, panels, and other not necessarily related information; and it is often difficult to automatically isolate this information. This is precisely the objective of…

信息检索 · 计算机科学 2012-10-24 Sergio López , Josep Silva , David Insa

The contextual information of Web images is investigated to address the issue of enriching their index characterizations with semantic descriptors and therefore bridge the semantic gap (i.e. the gap between the low-level content-based…

信息检索 · 计算机科学 2020-05-06 Fariza Fauzi , Mohammed Belkhatir

Due to the advancement in computer communication and storage technologies, large amount of image data is available on World Wide Web (WWW). In order to locate a particular set of images the available search engines may be used with the help…

信息检索 · 计算机科学 2024-09-05 R Rajkumar , M V Sudhamani

"Keyword Extraction" refers to the task of automatically identifying the most relevant and informative phrases in natural language text. As we are deluged with large amounts of text data in many different forms and content - emails, blogs,…

计算与语言 · 计算机科学 2019-08-22 Shibamouli Lahiri

Web content extraction - isolating a page's main content from surrounding boilerplate - is a prerequisite for search indexing, retrieval-augmented generation, NLP dataset construction, and large language model training. Progress in this…

计算与语言 · 计算机科学 2026-05-21 Murrough Foley

The contextual information of Web images is investigated to address the issue of characterizing their content with semantic descriptors and therefore bridge the semantic gap, i.e. the gap between their automated low-level representation in…

信息检索 · 计算机科学 2020-05-06 Fariza Fauzi , Mohammed Belkhatir

When searching the web, it is often possible that there are too many results available for ambiguous queries. Text snippets, extracted from the retrieved pages, are an indicator of the pages' usefulness to the query intention and can be…

信息检索 · 计算机科学 2009-03-24 N. Zotos , P. Tzekou , G. Tsatsaronis , L. Kozanidis , S. Stamou , I. Varlamis
‹ 上一页 1 2 3 10 下一页 ›