English
Related papers

Related papers: Longitudinal Sampling of URLs From the Wayback Mac…

200 papers

As of February, 2015, HTTP/2, the update to the 16-year-old HTTP 1.1, is officially complete. HTTP/2 aims to improve the Web experience by solving well-known problems (e.g., head of line blocking and redundant headers), while introducing…

Networking and Internet Architecture · Computer Science 2015-07-24 Matteo Varvello , Kyle Schomp , David Naylor , Jeremy Blackburn , Alessandro Finamore , Kostantina Papagiannaki

There are few studies that look closely at how the topology of the Internet evolves over time; most focus on snapshots taken at a particular point in time. In this paper, we investigate the evolution of the topology of the Autonomous…

Networking and Internet Architecture · Computer Science 2012-02-20 Benjamin Edwards , Steven Hofmeyr , George Stelle , Stephanie Forrest

The understanding of the immense and intricate topological structure of the World Wide Web (WWW) is a major scientific and technological challenge. This has been tackled recently by characterizing the properties of its representative graphs…

Networking and Internet Architecture · Computer Science 2008-01-23 M. Angeles Serrano , Ana Maguitman , Marian Boguna , Santo Fortunato , Alessandro Vespignani

Simple economic and performance arguments suggest appropriate lifetimes for main memory pages and suggest optimal page sizes. The fundamental tradeoffs are the prices and bandwidths of RAMs and disks. The analysis indicates that with…

Databases · Computer Science 2007-05-23 Jim Gray , Goetz Graefe

Longitudinal corpora like legal, corporate and newspaper archives are of immense value to a variety of users, and time as an important factor strongly influences their search behavior in these archives. While many systems have been…

Information Retrieval · Computer Science 2018-10-25 Jaspreet Singh , Avishek Anand

Nowadays, web archives preserve the history of large portions of the web. As medias are shifting from printed to digital editions, accessing these huge information sources is drawing increasingly more attention from national and…

Information Retrieval · Computer Science 2013-08-23 Zeynep Pehlivan , Benjamin Piwowarski , Stéphane Gançarski

World Wide Web consists of more than 50 billion pages online. It is highly dynamic i.e. the web continuously introduces new capabilities and attracts many people. Due to this explosion in size, the effective information retrieval system or…

Information Retrieval · Computer Science 2012-05-15 Sk. AbdulNabi , P. Premchand

Generative search engines increasingly determine whether online information is merely discoverable, cited as a source, or actually absorbed into generated answers. This paper proposes a two-stage measurement framework for Generative Engine…

Information Retrieval · Computer Science 2026-04-30 Zhang Kai , He Xinyue , Yao Jingang

Web scraping is a powerful technique that extracts data from websites, enabling automated data collection, enhancing data analysis capabilities, and minimizing manual data entry efforts. Existing methods, wrappers-based methods suffer from…

Computation and Language · Computer Science 2024-09-27 Wenhao Huang , Zhouhong Gu , Chenghao Peng , Zhixu Li , Jiaqing Liang , Yanghua Xiao , Liqian Wen , Zulong Chen

We conducted a preliminary field study to understand the current state of personal digital archiving in practice. Our aim is to design a service for the long-term storage, preservation, and access of digital belongings by examining how…

Digital Libraries · Computer Science 2007-05-23 Catherine C. Marshall , Sara Bly , Francoise Brun-Cottan

Websites employ third-party ads and tracking services leveraging cookies and JavaScript code, to deliver ads and track users' behavior, causing privacy concerns. To limit online tracking and block advertisements, several ad-blocking (black)…

Cryptography and Security · Computer Science 2019-06-04 Saad Sajid Hashmi , Muhammad Ikram , Mohamed Ali Kaafar

We analyze the online response to the preprint publication of a cohort of 4,606 scientific articles submitted to the preprint database arXiv.org between October 2010 and May 2011. We study three forms of responses to these preprints:…

Social and Information Networks · Computer Science 2015-06-04 Xin Shuai , Alberto Pepe , Johan Bollen

We present a framework for web-scale archiving of the dark web. While commonly associated with illicit and illegal activity, the dark web provides a way to privately access web information. This is a valuable and socially beneficial tool to…

Digital Libraries · Computer Science 2021-07-12 Justin F. Brunelle , Ryan Farley , Grant Atkins , Trevor Bostic , Marites Hendrix , Zak Zebrowski

Browser fingerprinting consists in collecting attributes from a web browser to build a browser fingerprint. In this work, we assess the adequacy of browser fingerprints as an authentication factor, on a dataset of 4,145,408 fingerprints…

Cryptography and Security · Computer Science 2021-06-23 Nampoina Andriamilanto , Tristan Allard , Gaëtan Le Guelvouit

Researchers have studied Internet censorship for nearly as long as attempts to censor contents have taken place. Most studies have however been limited to a short period of time and/or a few countries; the few exceptions have traded off…

Cryptography and Security · Computer Science 2019-07-11 Arian Akhavan Niaki , Shinyoung Cho , Zachary Weinberg , Nguyen Phong Hoang , Abbas Razaghpanah , Nicolas Christin , Phillipa Gill

Missing web pages, URIs that return the 404 "Page Not Found" error or the HTTP response code 200 but dereference unexpected content, are ubiquitous in today's browsing experience. We use Internet search engines to relocate such missing…

Information Retrieval · Computer Science 2010-04-19 Martin Klein , Jeffery Shipman , Michael L. Nelson

We performed a large-scale crawl of the World Wide Web, covering 6.9 Million domains and 57 Million subdomains, including all high-traffic sites of the Internet. We present a study of the correlations found between quantities measuring the…

Physics and Society · Physics 2015-06-12 G. A. Luduena , H. Meixner , Gregor Kaczor , Claudius Gros

The TREC 2009 web ad hoc and relevance feedback tasks used a new document collection, the ClueWeb09 dataset, which was crawled from the general Web in early 2009. This dataset contains 1 billion web pages, a substantial fraction of which…

Information Retrieval · Computer Science 2015-03-17 Gordon V. Cormack , Mark D. Smucker , Charles L. A. Clarke

Progress in AI is driven largely by the scale and quality of training data. Despite this, there is a deficit of empirical analysis examining the attributes of well-established datasets beyond text. In this work we conduct the largest and…

The Internet has significantly expanded the potential for global collaboration, allowing millions of users to contribute to collective projects like Wikipedia. While prior work has assessed the success of online collaborations, most…

Computers and Society · Computer Science 2025-03-17 Abraham Israeli , David Jurgens , Daniel Romero