English
Related papers

Related papers: A Brief History of Web Crawlers

200 papers

This articles surveys the existing literature on the methods currently used by web services to track the user online as well as their purposes, implications, and possible user's defenses. A significant majority of reviewed articles and web…

Computers and Society · Computer Science 2015-07-29 Tomasz Bujlow , Valentín Carela-Español , Josep Solé-Pareta , Pere Barlet-Ros

Indexing the Web is becoming a laborious task for search engines as the Web exponentially grows in size and distribution. Presently, the most effective known approach to overcome this problem is the use of focused crawlers. A focused…

Information Retrieval · Computer Science 2015-10-02 Ali Seyfi

Web traffic is a valuable data source, typically used in the marketing space to track brand awareness and advertising effectiveness. However, web traffic is also a rich source of information for cybersecurity monitoring efforts. To better…

Information Retrieval · Computer Science 2019-04-04 Han Qin , Kit Riehle , Haozhen Zhao

The random surfer model is a frequently used model for simulating user navigation behavior on the Web. Various algorithms, such as PageRank, are based on the assumption that the model represents a good approximation of users browsing a…

Social and Information Networks · Computer Science 2015-08-05 Florian Geigl , Daniel Lamprecht , Rainer Hofmann-Wellenhof , Simon Walk , Markus Strohmaier , Denis Helic

Web refresh crawling is the problem of keeping a cache of web pages fresh, that is, having the most recent copy available when a page is requested, given a limited bandwidth available to the crawler. Under the assumption that the change and…

Over the past three decades, computers have managed to make their way into a majority of households. Due to this enormous transition, the surge in the internets popularity was inevitable. Just like everything else, whatever has a pro also…

Cryptography and Security · Computer Science 2024-09-02 C. Amuthadevi , Sparsh Srivastava , Raghav Khatoria , Varun Sangwan

Web browsers are integral parts of everyone's daily life. They are commonly used for security-critical and privacy sensitive tasks, like banking transactions and checking medical records. Unfortunately, modern web browsers are too complex…

Cryptography and Security · Computer Science 2022-01-03 Jungwon Lim , Yonghwi Jin , Mansour Alharthi , Xiaokuan Zhang , Jinho Jung , Rajat Gupta , Kuilin Li , Daehee Jang , Taesoo Kim

A typical web search engine consists of three principal parts: crawling engine, indexing engine, and searching engine. The present work aims to optimize the performance of the crawling engine. The crawling engine finds new web pages and…

Networking and Internet Architecture · Computer Science 2012-01-20 Konstantin Avrachenkov , Alexander Dudin , Valentina Klimenok , Philippe Nain , Olga Semenova

Today World Wide Web (WWW) has become a huge ocean of information and it is growing in size everyday. Downloading even a fraction of this mammoth data is like sailing through a huge ocean and it is a challenging task indeed. In order to…

Information Retrieval · Computer Science 2011-02-04 Debajyoti Mukhopadhyay , Sajal Mukherjee , Soumya Ghosh , Saheli Kar , Young-Chon Kim

Understanding how people interact with the web is key for a variety of applications, e.g., from the design of effective web pages to the definition of successful online marketing campaigns. Browsing behavior has been traditionally…

Computers and Society · Computer Science 2021-05-05 Luca Vassio , Idilio Drago , Marco Mellia , Zied Ben Houidi , Mohamed Lamine Lamali

Websites are regarded as domains of limitless information which anyone and everyone can access. The new trend of technology put us to change the way we are doing our business. The Internet now is fastly becoming a new place for business and…

Information Retrieval · Computer Science 2021-09-03 Ikechukwu Onyenwe , Ebele Onyedinma , Chidinma Nwafor , Obinna Agbata

Log files contain information about User Name, IP Address, Time Stamp, Access Request, number of Bytes Transferred, Result Status, URL that Referred and User Agent. The log files are maintained by the web servers. By analysing these log…

Databases · Computer Science 2011-02-01 L. K. Joshila Grace , V. Maheswari , Dhinaharan Nagamalai

Searching accounts for one of the most frequently performed computations over the Internet as well as one of the most important applications of outsourced computing, producing results that critically affect users' decision-making behaviors.…

Soon after the invention of the Internet, the recommender system emerged and related technologies have been extensively studied and applied by both academia and industry. Currently, recommender system has become one of the most successful…

Information Retrieval · Computer Science 2022-09-07 Zhenhua Dong , Zhe Wang , Jun Xu , Ruiming Tang , Jirong Wen

This paper proposes a novel model for web crawling suitable for large-scale web data acquisition. This model first divides web data into several sub-data, with each sub-data corresponding to a thread task. In each thread task, web crawling…

Databases · Computer Science 2024-07-16 Weijie. Jiang

The World Wide Web's connectivity is greatly attributed to the HTTP protocol, with HTTP messages offering informative header fields that appeal to disciplines like web security and privacy, especially concerning web tracking. Despite…

Cryptography and Security · Computer Science 2025-02-28 Wolf Rieder , Philip Raschke , Thomas Cory

Given the vast scale of the Web, crawling prioritisation techniques based on link graph traversal, popularity, link analysis, and textual content are frequently applied to surface documents that are most likely to be valuable. While…

Information Retrieval · Computer Science 2025-07-03 Francesca Pezzuti , Sean MacAvaney , Nicola Tonellotto

Researchers in the Digital Humanities and journalists need to monitor, collect and analyze fresh online content regarding current events such as the Ebola outbreak or the Ukraine crisis on demand. However, existing focused crawling…

Digital Libraries · Computer Science 2017-07-31 Gerhard Gossen , Elena Demidova , Thomas Risse

Browser fingerprinting is a pervasive online tracking technique used increasingly often for profiling and targeted advertising. Prior research on the prevalence of fingerprinting heavily relied on automated web crawls, which inherently…

Cryptography and Security · Computer Science 2025-02-04 Meenatchi Sundaram Muthu Selva Annamalai , Igor Bilogrevic , Emiliano De Cristofaro

We consider a task of scheduling a crawler to retrieve content from several sites with ephemeral content. A user typically loses interest in ephemeral content, like news or posts at social network groups, after several days or hours. Thus,…

Information Retrieval · Computer Science 2015-03-31 Konstantin Avrachenkov , Vivek Borkar