English
Related papers

Related papers: Estimating Absolute Web Crawl Coverage From Longit…

200 papers

Random walks have been proposed as a simple method of efficiently searching, or disseminating information throughout, communication and sensor networks. In nature, animals (such as ants) tend to follow correlated random walks, i.e., random…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-02-03 Graeme Smith , J. W. Sanders , Qin Li

In the course of web research it is often necessary to estimate the creation datetime for web resources (in the general case, this value can only be estimated). While it is feasible to manually establish likely datetime values for small…

Information Retrieval · Computer Science 2013-04-19 Hany M. SalahEldeen , Michael L. Nelson

Prior work on web archive profiling were focused on Archival Holdings to describe what is present in an archive. This work defines and explores Archival Voids to establish a means to represent portions of URI spaces that are not present in…

Digital Libraries · Computer Science 2021-08-10 Sawood Alam , Michele C. Weigle , Michael L. Nelson

We consider a task of scheduling a crawler to retrieve content from several sites with ephemeral content. A user typically loses interest in ephemeral content, like news or posts at social network groups, after several days or hours. Thus,…

Information Retrieval · Computer Science 2015-03-31 Konstantin Avrachenkov , Vivek Borkar

The design of effective online caching policies is an increasingly important problem for content distribution networks, online social networks and edge computing services, among other areas. This paper proposes a new algorithmic toolbox for…

Networking and Internet Architecture · Computer Science 2022-09-28 Naram Mhaisen , George Iosifidis , Douglas Leith

Web archives preserve unique and historically valuable information. They hold a record of past events and memories published by all kinds of people, such as journalists, politicians and ordinary people who have shared their testimony and…

Digital Libraries · Computer Science 2021-08-04 Miguel Costa , Julien Masanès

The Web graph is a giant social network whose properties have been measured and modeled extensively in recent years. Most such studies concentrate on the graph structure alone, and do not consider textual properties of the nodes.…

Information Retrieval · Computer Science 2018-02-15 Soumen Chakrabarti , Mukul M. Joshi , Kunal Punera , David M. Pennock

Given the vast scale of the Web, crawling prioritisation techniques based on link graph traversal, popularity, link analysis, and textual content are frequently applied to surface documents that are most likely to be valuable. While…

Information Retrieval · Computer Science 2025-07-03 Francesca Pezzuti , Sean MacAvaney , Nicola Tonellotto

Users' detailed browsing activity - such as what sites they are spending time on and for how long, and what tabs they have open and which one is focused at any given time - is useful for a number of research and practical applications.…

Human-Computer Interaction · Computer Science 2021-02-09 Geza Kovacs

In this paper we present a preliminary analysis over the largest publicly accessible web dataset: the Common Crawl Corpus. We measure nine web characteristics from two levels of granularity using MapReduce and we comment on the initial…

Information Retrieval · Computer Science 2014-09-30 Vasilis Kolias , Ioannis Anagnostopoulos , Eleftherios Kayafas

Large language models (LLMs) rely heavily on web-scale datasets like Common Crawl, which provides over 80\% of training data for some modern models. However, the indiscriminate nature of web crawling raises challenges in data quality,…

Computation and Language · Computer Science 2025-09-01 Inés Altemir Marinas , Anastasiia Kucherenko , Andrei Kucharavy

A large number of URLs are made public by various platforms for security analysis, archiving, and paste sharing -- such as VirusTotal, URLScan.io, Hybrid Analysis, the Wayback Machine, and RedHunt. These services may unintentionally expose…

Cryptography and Security · Computer Science 2026-02-26 Tarek Ramadan , AbdelRahman Abdou , Mohammad Mannan , Amr Youssef

Active perception approaches select future viewpoints by using some estimate of the information gain. An inaccurate estimate can be detrimental in critical situations, e.g., locating a person in distress. However the true information gained…

Robotics · Computer Science 2026-04-17 Siming He , Yuezhan Tao , Igor Spasojevic , Vijay Kumar , Pratik Chaudhari

Nowadays, the huge amount of information distributed through the Web motivates studying techniques to be adopted in order to extract relevant data in an efficient and reliable way. Both academia and enterprises developed several approaches…

Artificial Intelligence · Computer Science 2013-06-06 Emilio Ferrara , Robert Baumgartner

Big data streams are grasping increasing attention with the development of modern science and information technology. Due to the incompatibility of limited computer memory to high volume of streaming data, real-time methods without…

Methodology · Statistics 2023-06-29 Chunbai Tao , Shanshan Wang

This paper presents a comprehensive analysis of global web usage patterns based on data from SimilarWeb, a leading source for estimating web traffic. Leveraging a dataset comprising over 250,000 websites, we estimate the total web traffic…

Computers and Society · Computer Science 2024-11-27 Henrique S. Xavier

Commercial web search engines employ near-duplicate detection to ensure that users see each relevant result only once, albeit the underlying web crawls typically include (near-)duplicates of many web pages. We revisit the risks and…

With the large-scale penetration of the internet, for the first time, humanity has become linked by a single, open, communications platform. Harnessing this fact, we report insights arising from a unified internet activity and location…

Economics · Quantitative Finance 2017-01-23 Klaus Ackermann , Simon D Angus , Paul A Raschky

Online Social Network (OSN) is one of the most hottest services in the past years. It preserves the life of users and provides great potential for journalists, sociologists and business analysts. Crawling data from social network is a basic…

Social and Information Networks · Computer Science 2013-12-10 Rui Guo , Hongzhi Wang , Mengwen Chen , Jianzhong Li , Hong Gao

The ability to automatically detect fraudulent escrow websites is important in order to alleviate online auction fraud. Despite research on related topics, fake escrow website categorization has received little attention. In this study we…

Computers and Society · Computer Science 2013-09-30 Ahmed Abbasi , Hsinchun Chen
‹ Prev 1 4 5 6 7 8 10 Next ›