English
Related papers

Related papers: Longitudinal Sampling of URLs From the Wayback Mac…

200 papers

Recent advances of preservation technologies have led to an increasing number of Web archive systems and collections. These collections are valuable to explore the past of the Web, but their value can only be uncovered with effective access…

Information Retrieval · Computer Science 2017-01-17 Khoi Duy Vo , Tuan Tran , Tu Ngoc Nguyen , Xiaofei Zhu , Wolfgang Nejdl

Significant parts of cultural heritage are produced on the web during the last decades. While easy accessibility to the current web is a good baseline, optimal access to the past web faces several challenges. This includes dealing with…

Digital Libraries · Computer Science 2017-01-31 Nattiya Kanhabua , Philipp Kemkes , Wolfgang Nejdl , Tu Ngoc Nguyen , Felipe Reis , Nam Khanh Tran

Over the past two decades, a desire to reduce transit cost, improve control over routing and performance, and enhance the quality of experience for users, has yielded a more densely connected, flat network with fewer hops between sources…

Networking and Internet Architecture · Computer Science 2023-03-07 Esteban Carisimo , Mia Weaver , Paul Barford , Fabián E. Bustamante

Over the last 30 years, the World Wide Web has changed significantly. In this paper, we argue that common practices to prepare web pages for delivery conflict with many efforts to present content with minimal latency, one fundamental goal…

Networking and Internet Architecture · Computer Science 2024-03-26 Lucas Vogel , Thomas Springer , Matthias Wählisch

The World Wide Web is the most wide known information source that is easily available and searchable. It consists of billions of interconnected documents Web pages are authored by millions of people. Accesses made by various users to pages…

Databases · Computer Science 2014-08-26 Priyanka Verma , Nishtha Kesswani

Curated web archive collections contain focused digital contents which are collected by archiving organizations to provide a representative sample covering specific topics and events to preserve them for future exploration and analysis. In…

Digital Libraries · Computer Science 2017-02-02 Zeon Trevor Fernando , Ivana Marenzi , Wolfgang Nejdl , Rishita Kalyani

Although web advertisements represent an inimitable part of digital cultural heritage, serious archiving and replay challenges persist. To explore these challenges, we created a dataset of 279 archived ads. We encountered five problems in…

Digital Libraries · Computer Science 2025-09-24 Travis Reid , Alex H. Poole , Hyung Wook Choi , Christopher Rauch , Mat Kelly , Michael L. Nelson , Michele C. Weigle

Recently, reproducibility has become a cornerstone in the security and privacy research community, including artifact evaluations and even a new symposium topic. However, Web measurements lack tools that can be reused across many…

Cryptography and Security · Computer Science 2025-06-05 Florian Hantke , Peter Snyder , Hamed Haddadi , Ben Stock

With social media datasets being increasingly shared by researchers, it also presents the caveat that those datasets are not always completely replicable. Having to adhere to requirements of platforms like Twitter, researchers cannot…

Digital Libraries · Computer Science 2018-03-08 Arkaitz Zubiaga

Purpose: Advanced usage of Web Analytics tools allows to capture the content of user queries. Despite their relevant nature, the manual analysis of large volumes of user queries is problematic. This paper demonstrates the potential of using…

Information Retrieval · Computer Science 2017-09-25 Anne Chardonnens , Ettore Rizza , Mathias Coeckelbergs , Seth van Hooland

The increasing adoption of econometric and machine-learning approaches by empirical researchers has led to a widespread use of one data collection method: web scraping. Web scraping refers to the use of automated computer programs to access…

General Economics · Economics 2023-08-07 Jens Foerderer

The use of citation counts to assess the impact of research articles is well established. However, the citation impact of an article can only be measured several years after it has been published. As research articles are increasingly…

Information Retrieval · Computer Science 2007-05-23 Tim Brody , Stevan Harnad

Webpages change over time, and web archives hold copies of historical versions of webpages. Users of web archives, such as journalists, want to find and view changes on webpages over time. However, the current search interfaces for web…

Information Retrieval · Computer Science 2023-05-02 Lesley Frew , Michael L. Nelson , Michele C. Weigle

Research articles are being shared in increasing numbers on multiple online platforms. Although the scholarly impact of these articles has been widely studied, the online interest determined by how long the research articles are shared…

Digital Libraries · Computer Science 2022-09-15 Murtuza Shahzad , Hamed Alhoori , Reva Freedman , Shaikh Abdul Rahman

This paper maps the national UK web presence on the basis of an analysis of the .uk domain from 1996 to 2010. It reviews previous attempts to use web archives to understand national web domains and describes the dataset. Next, it presents…

Digital Libraries · Computer Science 2023-01-05 Scott A. Hale , Taha Yasseri , Josh Cowls , Eric T. Meyer , Ralph Schroeder , Helen Margetts

Many web sites are transitioning how they construct their pages. The conventional model is where the content is embedded server-side in the HTML and returned to the client in an HTTP response. Increasingly, sites are moving to a model where…

Digital Libraries · Computer Science 2023-05-03 Michele C. Weigle , Michael L. Nelson , Sawood Alam , Mark Graham

A broad range of research areas including Internet measurement, privacy, and network security rely on lists of target domains to be analysed; researchers make use of target lists for reasons of necessity or efficiency. The popular Alexa…

Networking and Internet Architecture · Computer Science 2018-09-25 Quirin Scheitle , Oliver Hohlfeld , Julien Gamba , Jonas Jelten , Torsten Zimmermann , Stephen D. Strowes , Narseo Vallina-Rodriguez

Humans and large language models (LLMs) now co-produce and co-consume the web's shared knowledge archives. Such human-AI collective knowledge ecosystems contain feedback loops with both benefits (e.g., faster growth, easier learning) and…

Computers and Society · Computer Science 2026-01-29 Buddhika Nettasinghe , Kang Zhao

Most archived HTML pages embed other web resources, such as images and stylesheets. Playback of the archived web pages typically provides only the capture date (or Memento-Datetime) of the root resource and not the Memento-Datetime of the…

Digital Libraries · Computer Science 2014-10-07 Scott G. Ainsworth , Michael L. Nelson , Herbert Van de Sompel

We present 3DLNews, a novel dataset with local news articles from the United States spanning the period from 1996 to 2024. It contains almost 1 million URLs (with HTML text) from over 14,000 local newspapers, TV, and radio stations across…

Information Retrieval · Computer Science 2024-08-12 Gangani Ariyarathne , Alexander C. Nwala