中文
相关论文

相关论文: The Impact of Main Content Extraction on Near-Dupl…

200 篇论文

Nowadays, digital content is widespread and simply redistributable, either lawfully or unlawfully. For example, after images are posted on the internet, other web users can modify them and then repost their versions, thereby generating…

计算机视觉与模式识别 · 计算机科学 2020-09-08 K. K. Thyagharajan , G. Kalaiarasi

Various software features such as classes, methods, requirements, and tests often have similar functionality. This can lead to emergence of duplicates in their descriptive documentation. Uncontrolled duplicates created via copy/paste hinder…

A focused crawler traverses the web selecting out relevant pages to a predefined topic and neglecting those out of concern. While surfing the internet it is difficult to deal with irrelevant pages and to predict which links lead to quality…

信息检索 · 计算机科学 2009-06-30 Anshika Pal , Deepak Singh Tomar , S. C. Shrivastava

Template extraction is the process of isolating the template of a given webpage. It is widely used in several disciplines, including webpages development, content extraction, block detection, and webpages indexing. One of the main goals of…

信息检索 · 计算机科学 2014-09-10 Julián Alarte , David Insa , Josep Silva , Salvador Tamarit

Template detection and content extraction are two of the main areas of information retrieval applied to the Web. They perform different analyses over the structure and content of webpages to extract some part of the document. However, their…

信息检索 · 计算机科学 2022-07-19 Julián Alarte , Josep Silva

The importance of an efficient and scalable document similarity detection system is undeniable nowadays. Search engines need batch text similarity measures to detect duplicated and near-duplicated web pages in their indexes in order to…

信息检索 · 计算机科学 2018-10-09 Hamid Mohammadi , Amin Nikoukaran

Detecting near duplicate images is fundamental to the content ecosystem of photo sharing web applications. However, such a task is challenging when involving a web-scale image corpus containing billions of images. In this paper, we present…

计算机视觉与模式识别 · 计算机科学 2022-09-20 Andrey Gusev , Jiajing Xu

In this paper we propose a bayesian approach for near-duplicate image detection, and investigate how different probabilistic models affect the performance obtained. The task of identifying an image whose metadata are missing is often…

计算机视觉与模式识别 · 计算机科学 2021-08-23 Lucas Moutinho Bueno , Eduardo Valle , Ricardo da Silva Torres

The massive spread of visual content through the web and social media poses both challenges and opportunities. Tracking visually-similar content is an important task for studying and analyzing social phenomena related to the spread of such…

信息检索 · 计算机科学 2022-03-15 Hana Matatov , Mor Naaman , Ofra Amir

Contemporary software documentation is as complicated as the software itself. During its lifecycle, the documentation accumulates a lot of near duplicate fragments, i.e. chunks of text that were copied from a single source and were later…

软件工程 · 计算机科学 2018-10-10 D. V. Luciv , D. V. Koznov , G. A. Chernishev , A. N. Terekhov

Completeness of a knowledge graph is an important quality dimension and factor on how well an application that makes use of it performs. Completeness can be improved by performing knowledge enrichment. Duplicate detection aims to find…

数据库 · 计算机科学 2022-07-21 Juliette Opdenplatz , Umutcan Şimşek , Dieter Fensel

We describe a system that helps identify manuscripts submitted to multiple journals at the same time. Also, we discuss potential applications of the near-duplicate detection technology when run with manuscript text content, including…

In the context of End-to-End testing of web applications, automated exploration techniques (a.k.a. crawling) are widely used to infer state-based models of the site under test. These models, in which states represent features of the web…

软件工程 · 计算机科学 2021-08-31 Anna Corazza , Sergio Di Martino , Adriano Peron , Luigi Libero Lucio Starace

In large-scale data analysis, near-duplicates are often a problem. For example, with two near-duplicate phishing emails, a difference in the salutation (Mr versus Ms) is not essential, but whether it is bank A or B is important. The…

密码学与安全 · 计算机科学 2023-08-24 Pieter Hartel , Eljo Haspels , Mark van Staalduinen , Octavio Texeira

Web templates are one of the main development resources for website engineers. Templates allow them to increase productivity by plugin content into already formatted and prepared pagelets. For the final user templates are also useful,…

信息检索 · 计算机科学 2015-01-12 Julián Alarte , David Insa , Josep Silva , Salvador Tamarit

One of the important factors that make a search engine fast and accurate is a concise and duplicate free index. In order to remove duplicate and near-duplicate documents from the index, a search engine needs a swift and reliable duplicate…

信息检索 · 计算机科学 2019-09-26 Hamid Mohammadi , Seyed Hossein Khasteh

Semantic web is a web of future. The Resource Description Framework (RDF) is a language to represent resources in the World Wide Web. When these resources are queried the problem of duplicate query results occurs. The present techniques…

数据库 · 计算机科学 2013-05-14 1Oumair Naseer , 2Ayesha Naseer , 3Atif Ali Khan , 4Humza Naseer

There is an explosive growth of information in the World Wide Web thus posing a challenge to Web users to extract essential knowledge from the Web. Search engines help us to narrow down the search in the form of Search Engine Result Pages…

信息检索 · 计算机科学 2013-03-26 Srikantaiah K C , Suraj M , Venugopal K R , L M Patnaik

Search engines have become an indispensable tool for browsing information on the Internet. The user, however, is often annoyed by redundant results from irrelevant Web pages. One reason is because search engines also look at non-informative…

信息检索 · 计算机科学 2019-11-27 Dat Quoc Nguyen , Dai Quoc Nguyen , Son Bao Pham , The Duy Bui

Job descriptions are posted on many online channels, including company websites, job boards or social media platforms. These descriptions are usually published with varying text for the same job, due to the requirements of each platform or…

计算与语言 · 计算机科学 2024-06-11 Matthias Engelbach , Dennis Klau , Maximilien Kintz , Alexander Ulrich
‹ 上一页 1 2 3 10 下一页 ›